UNAM Nullifies Thousands of Exams After AI Use Distorts Mexico's Largest Admissions Test
AI & ML

UNAM Nullifies Thousands of Exams After AI Use Distorts Mexico's Largest Admissions Test

Mexico's National Autonomous University found perfect scores nearly quintupled on its first online entrance exam, and it is now forcing a pen and paper retest for tens of thousands of applicants.

PublishedAugust 3, 2026
Read time6 min read
Share

A record-scale exam meets a record-scale problem

The National Autonomous University of Mexico, UNAM, is the largest university in Latin America, and this spring it ran its first fully online undergraduate entrance exam in its history. Over three weekends spanning May and June, 158,727 applicants sat for the test from home rather than in a supervised hall, a scale that made traditional in-person proctoring impossible and put the university's newer remote integrity systems under a real stress test for the first time. The decision to go fully online had been framed internally as a modernization step, an overdue upgrade for an institution that processes admissions at a scale few universities anywhere in the world attempt in a single cycle.

Results went out on July 17, and within days UNAM's admissions office and a technical committee formed on July 27 were fielding a pattern in the data too large to dismiss as ordinary, isolated cheating. On August 2, Rector Leonardo Lomeli confirmed the university would nullify a portion of the exams and require an entirely new round of retesting, calling it one of the largest academic integrity investigations in UNAM's history and a direct consequence of moving a high stakes exam online before its detection tooling was ready for the scale of the attempt.

The number that gave it away

What exposed the problem was not a single flagged test session, it was the distribution of scores across the entire applicant pool once the technical committee sat down with the full data set. Perfect scores had averaged 3.5 percent of exams across the five prior admissions cycles, from 2021 through 2025, a stable baseline the university had used for years to sanity check results. In 2026 that figure jumped to 16.3 percent, a nearly fivefold increase in a single testing window that no ordinary shift in applicant preparation, question difficulty, or study habits could plausibly explain on its own.

That kind of anomaly is exactly what population level analysis is built to catch, and exactly what per-session proctoring tools are structurally unable to see, because each of those tools is only ever looking at one test taker at a time. UNAM's technical committee treated the score distribution itself as the primary piece of evidence, then worked backward to build individual cases against roughly 3,175 exams, about 2 percent of the total, once the statistical signal told investigators precisely where in the applicant pool to concentrate their review.

Where the proctoring layer actually helped, and where it did not

UNAM's online proctoring system did its job at the individual level, and it is worth crediting what it caught before turning to what it missed. It flagged phone usage during the test window, coaching from other people present in the room, and cases of apparent identity substitution, where someone other than the registered applicant appears to have sat the exam on their behalf. Those are the classic failure modes remote proctoring vendors design their products around, the ones that show up clearly on camera or in browser telemetry, and on those the system performed exactly as intended.

What it could not do was catch AI assisted answering, because a strong answer produced with help from a chatbot looks, session by session, indistinguishable from a strong answer a well prepared applicant produced on their own. A test taker quietly consulting an AI tool off screen does not trip a webcam alert or a browser lockdown rule, and no individual session looks anomalous in isolation. The tell only became visible once someone looked at the aggregate result across the full applicant pool, a detection gap every proctoring vendor selling into high stakes assessment now has to reckon with.

The cost of getting it right after the fact

Fixing this after results were already public was expensive in ways that go well beyond the logistics of a retest. UNAM suspended undergraduate enrollment procedures for the entire 2026 to 2027 admissions cycle while the investigation ran its course, an unusual and disruptive step for an institution processing applications at this scale, and one that pushed uncertainty onto every applicant in the pool, not just the ones under specific suspicion. The rector's public apology, offered directly to applicants who had already been preliminarily admitted, put a human face on a process the university itself had to unwind after telling those students they had gotten in.

The retest population is not limited to the roughly 3,175 exams UNAM nullified outright. The university is also requiring new tests from rejected applicants whose scores matched or exceeded the lowest passing threshold recorded across the 2021 through 2025 cycles, on the logic that a distorted curve may have pushed some legitimately capable applicants below the cut line while inflated scores pushed others above it. That decision widens the retest population, and the disruption, considerably beyond the cases where UNAM already has the clearest evidence of wrongdoing.

Back to pen and paper is itself the headline

UNAM confirmed the makeup exam will be administered in person and on paper, though it has not yet announced specific dates or locations for the retest. For a university that had just made its first attempt at a fully remote, at scale entrance exam, reverting all the way back to the oldest anti-cheating technology available, a proctored room and a printed booklet, is a candid admission that the remote version could not be trusted at this scale, at least not with the tools UNAM had in place this cycle.

That reversal matters more than any single number in this story, because it signals what a resourced, motivated institution concluded was actually necessary once the stakes became clear. Even with a dedicated technical committee, commercial proctoring software, and a rector willing to absorb the reputational hit of a public apology, UNAM decided the safest path forward was to remove the AI assisted attack surface entirely rather than attempt to patch around it with better monitoring.

What this means beyond admissions

UNAM's exam is an extreme case in scale, with nearly 160,000 applicants in a single cycle, but the underlying exposure is not unique to universities or to admissions testing specifically. Any organization running high stakes assessment at scale through a remote channel, professional certification bodies, licensing boards, corporate skills validation programs, internal promotion exams, carries the same structural gap: proctoring tools tuned to catch individual violations on camera while missing a systemic shift in the overall score distribution that only becomes visible in aggregate.

The practical lesson for technology leaders who own certification, licensing, or hiring assessment programs is to build population level score monitoring into the program from the outset, rather than as an afterthought triggered by a scandal after the fact. Compare each cohort's score distribution against historical baselines on a rolling basis, and treat a statistically implausible jump in top scores or perfect results as a detection event in its own right, one that gets investigated regardless of what the proctoring vendor's session level alerts happen to report.

Tagged#news#edtech#education#learning#lms#ai-education#unam#mexico-admissions-exam#academic-integrity-crisis#exam-invalidation#higher-education-mexico