인공지능이 16만 명의 학생들의 부정행위를 막을 수 없었던 이유

작성자

카테고리:

← 피드로
DEV Community · Mohit Geryani · 2026-08-07 개발(SW)

Every AI security system is built on a simple assumption:

If you can observe enough behavior, you can detect bad behavior.

That’s the idea behind AI-Based exam proctoring.

Watch students through their webcams. Lock down their browsers. Spot suspicious movements. Alert human supervisors when something looks wrong.

It feels like a robust security plan but earlier this year, one of the largest AI-proctored exams ever conducted proved that assumption wasn’t nearly as reliable as it sounded.

Nearly 160,000 applicants sat for the entrance exam to Mexico’s largest university, UNAM, under AI surveillance. The system monitored webcams, restricted computers with a lockdown browser, and even routed suspicious activity to human reviewers.

Despite all of that, something was clearly wrong.

When the results arrived, the number of top-scoring students had increased by almost five times compared to previous years. The anomaly was so significant that the university eventually decided around 58,000 applicants would have to retake the exam in person, just because a system specifically designed to prevent cheating failed at scale.

AI Didn’t Watch the Answers. It Watched the Students.

Unlike ChatGPT or Claude, exam proctoring AI isn’t designed to understand what a student is writing. Its job is to watch how they’re taking the exam.

For the UNAM entrance test, applicants were required to install a lockdown browser that blocked common ways of accessing outside information. Opening new tabs, switching applications, copying text, or searching the web was heavily restricted.

At the same time, an AI-powered webcam monitoring system continuously analyzed each candidate throughout the exam.

The software looked for behaviors it considered suspicious like Someone else appearing in front of the camera, Student leaving the frame, Looking away repeatedly, Using a phone or wearing hidden earphones and other unusual movements that might suggest outside assistance.

Instead of replacing human invigilators entirely, the system generated alerts that were reviewed by supervisors. Roughly one human monitored every 150 candidates, stepping in only when the AI flagged something unusual.

It seemed like a layered defense.

Lock the computer -> Watch the student -> Escalate suspicious behavior to a human.

It sounds robust but every one of those layers shared the same blind spot.

The Blind Spot Was Never the Webcam

Most remote proctoring systems are built around a simple idea: if you can watch a student closely enough, you can stop them from cheating.

But modern AI has changed what cheating looks like.

A student no longer needs to pull out a phone in front of the camera or glance at handwritten notes taped to the wall. They can position a second monitor just outside the webcam’s field of view. They can receive answers through an earpiece hidden beneath their hair. They can even ask another AI model for help on a separate device the proctoring software never sees.

The system can detect movement.

It can detect faces.

It can even detect when someone leaves the frame.

What it cannot detect is everything happening outside the frame.

According to reports following the exam, online communities were already sharing tips before the test began. Students discussed placing additional screens beyond the webcam’s view, hiding wireless earbuds, and finding ways to work around the monitoring software without triggering alerts.

If even a fraction of those methods worked, the consequences become obvious.

The AI wasn’t watching the entire testing environment.

It was only watching what its camera could see.

When the Numbers Stopped Making Sense

If the university had only caught a handful of students cheating, this story probably wouldn’t exist.

What forced UNAM to investigate was the data.

Between 2021 and 2025, only 3.5% of applicants scored 100 or more out of 120 on the university’s entrance exam.

This year, that number jumped to 16.3%.

The pattern became even harder to ignore at the very top of the scoreboard.

Historically, fewer than 1% of applicants scored 110 or higher.

In 2026, that figure climbed to 5.5%—an increase that statistical models simply couldn’t explain as a stronger group of students.

Something had changed.

The university formed an independent commission to review the entire admission process, from the online testing platform to the proctoring system itself. After examining the results, the commission concluded that the integrity of the exam could no longer be guaranteed.

Its recommendation was drastic.

Around 58,000 applicants including students who had already earned admission would have to sit for another exam and this time in person.

This Was Never Just an Exam Problem

It’s tempting to treat this as a story about students finding creative ways to cheat.

But that’s only part of what happened.

The bigger lesson is that AI surveillance has limits.

An AI model can detect whether someone leaves their seat. It can estimate where a person is looking. It can flag unusual movement patterns far faster than any human supervisor.

What it cannot do is understand intent.

A student staring at a second monitor just outside the webcam’s view may look perfectly normal. Someone quietly listening through a hidden earpiece may never trigger an alert. Another person could simply exploit a weakness no one anticipated when the system was designed.

The more sophisticated people become, the more surveillance systems end up reacting to yesterday’s tactics instead of tomorrow’s ones.

That’s not unique to online exams.

Banks use AI to detect fraud.

Companies use AI to monitor employees.

Governments use AI to identify suspicious behavior.

In every case, the same limitation applies.

AI can only reason about the information it’s allowed to observe.

The Future of AI Proctoring Can’t Be More Surveillance

UNAM’s decision to require nearly 58,000 students to retake the exam was an admission that the university could no longer trust the system designed to stop it.

Ironically, the solution wasn’t deploying a more advanced AI model or adding another layer of software.

It was bringing students back into a physical examination hall.

Sometimes, the answer isn’t building a smarter surveillance system.

It’s redesigning the process so surveillance matters less in the first place.

As AI becomes more capable, we’ll continue using it to monitor workplaces, verify identities, prevent fraud, and protect high-stakes systems.

But this incident is a reminder that surveillance alone is rarely enough.

The moment people understand what an AI system can see, they start looking for what it can’t.

And that’s a challenge every AI-powered security system, not just exam proctoring will have to solve.

원문에서 계속 ↗

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다