What happened
An AI exam proctoring rollout at UNAM, Mexico's largest university, has collapsed into one of the biggest testing scandals in recent memory. Nearly 160,000 applicants sat the entrance exam remotely for the first time this year, using Respondus LockDown Browser to block outside apps and a webcam monitoring system from Territorium that relied on AI algorithms to flag suspicious behavior. The exam ran over several weeks from late May through early June 2026.
The results didn't add up. Between 2021 and 2025, only 3.5 percent of test takers scored 100 or more out of 120 questions. This year that number jumped to 16.3 percent. At the very top of the scale, the gap was even starker: scores of 110-plus went from 0.9 percent of applicants historically to 5.5 percent this cycle. That kind of jump doesn't happen by accident, and it triggered immediate accusations of mass cheating.
UNAM convened an expert commission to investigate, and its conclusion was blunt: hold a mandatory in-person "control exam." The retake won't just apply to this year's applicants — it covers everyone who would have qualified for admission based on minimum passing scores dating back to 2021, roughly 58,000 people in total. UNAM's rector publicly apologized to honest applicants who did nothing wrong but must now prepare for and sit a second exam, calling the control test "necessary to give certainty and guarantee equity in access." With classes originally set to begin August 10, the university is now racing against its own calendar.
Why it matters
This isn't just a Mexico City story — it's a warning shot for every institution betting on AI proctoring to secure high-stakes remote testing. The exam was multiple choice, not essay-based, which made it nearly impossible to spot the usual AI-cheating fingerprints, like a fully formed answer pasted straight into a text box. Investigators instead suspect a mix of old and new tricks: leaked questions, physical cheat sheets, earphones hidden under hair, monitors positioned just outside the webcam frame so students could consult ChatGPT off-camera, and even proxy test-takers sitting in for applicants entirely.
A New York Times report found that tips for beating the AI proctor were circulating widely online before the exam even opened. That's the core problem: AI webcam monitoring is built to catch a narrow set of behaviors — a second face in frame, a phone appearing in view — but it has no way to know what's happening just outside the camera's field of view. Lockdown browsers stop you from switching tabs on the same device; they do nothing about a second screen sitting a few feet away.
How to use it today
For educators, HR teams, and founders building assessment products, the UNAM case is a practical case study in where AI proctoring alone falls short — and where it still adds real value. AI monitoring is genuinely useful for flagging anomalies at scale across 160,000 test sessions; no human review team could watch that many webcams live. The mistake was treating AI flags as sufficient on their own rather than as a trigger for deeper human review, randomized question banks, or in-person spot checks.
If you're designing remote quizzes, certifications, or hiring assessments, a few things from this incident are worth building into your process immediately: randomize question order and pull from larger item pools so leaked answer keys lose value fast, layer in identity verification beyond just facial recognition, and keep a human-in-the-loop review for any score that lands in a statistically unusual range. If you need to quickly generate varied question sets, proofread assessment content, or prototype quiz logic without a big engineering lift, tools like the free AI utilities at [mykreatool.com](https://mykreatool.com) can help small teams and educators put together first drafts and randomized variants faster, which reduces reliance on a static, leak-prone question bank.
Who benefits
Oddly enough, the fallout creates opportunity for a few groups. Ed-tech vendors that combine AI proctoring with stronger identity verification and live human proctors stand to gain credibility as universities reassess vendor contracts. Testing-security consultants and forensic statisticians — the kind of experts UNAM had to bring in after the fact — are likely to see more demand for pre-launch audits rather than post-scandal cleanup. And honest students, frustrating as a retake is, ultimately benefit from a system that catches and corrects for mass cheating rather than letting inflated scores stand and displace them from spots they earned fairly.
Risks
The risks here cut in multiple directions. For UNAM, there's reputational damage and a massive logistics problem: reorganizing an in-person exam for 58,000 people in a matter of weeks, all while trying to keep the academic calendar intact. For applicants, it means added stress, travel costs, and lost time preparing for a test they thought they'd already passed — even though the commission has been clear that most of them did nothing wrong.
More broadly, this case highlights a structural risk in AI exam proctoring: false negatives at scale are much harder to detect than false positives. A single AI system incorrectly flagging one honest student draws immediate complaints. Thousands of students quietly bypassing the same system produces no complaints at all — just a skewed score distribution that only shows up after the fact, once the damage to trust is already done. Any institution rolling out AI-based remote testing should assume its proctoring software will eventually be probed and shared as a workaround online, and plan verification layers accordingly.
Conclusion
UNAM's AI exam proctoring failure shows that automated webcam monitoring and lockdown browsers, while useful, are not a complete substitute for layered verification, randomized content, and human oversight in high-stakes testing. With 58,000 students now facing a mandatory retake and classes on the line, the episode is becoming a reference case for how not to deploy AI proctoring at scale — and a reminder that any institution adopting similar tools should pressure-test them before, not after, the exam goes live.



Comments 0