What does AI flagging mean on a proctored exam?
AI flagging is the published middle layer of online proctoring: while your exam runs, the proctoring software watches the session and marks moments that look like they break the exam's rules, and a human being later reviews those marked moments and makes the actual decision. Pearson VUE states the model at booking - candidates consent to monitoring by human proctors and assistive AI tools throughout the test. Honorlock states it in one sentence: the software monitors the exam session and alerts a live test proctor if it detects any problems.
A flag is not an accusation, and nowhere in this source set is it a verdict. Proctorio publishes the clearest version: the software does not perform any type of algorithmic decision-making, such as determining whether a breach of exam integrity has occurred, and all integrity decisions are left to the exam administrator or institution. The published flag categories are ordinary things - background noise, looking away from the screen, leaving the camera view, a voice, an unauthorized application or browser tab.
This page covers the published picture in general terms: the model every vendor describes, the flag categories they name, what the software is published to watch for, who reviews flags and when, what the test taker sees after a review, and what happens next. Every claim comes from official vendor pages. Nothing here describes how any flag is computed, what any threshold is, or how any system might be evaded - no vendor publishes those details, and this page does not speculate about them.
The published model: software marks, a person decides
Read the candidate-facing pages of the major systems and one architecture appears everywhere. Pearson VUE's general OnVUE page lists what online testing requires: run a system test, find a distraction-free space, consent to monitoring by human proctors and assistive AI tools throughout your test, and observe the online exam requirements. The two layers are named in the same breath, and agreeing to both is a precondition of sitting the exam at all.
Honorlock describes the same split from the student side: the remote proctoring services combine the benefits of AI software with those of live test proctors, the software monitors the exam session, and it alerts a live test proctor if it detects any problems. The company highlights the consequence for the candidate - this means you won't be watched during the entire exam. The software is the always-on layer; the human is the on-call layer.
Proctortrack sells the same design as a product tier. Its ProctorLive AI level is published as exception-based live proctoring: a real-time hybrid model that couples live remote human proctors with AI-enhanced auto-proctoring intervention capabilities in cases of suspicious behaviors, cheating, or aid to a student, with integrity results analyzed further by AI. Whatever the branding, the shape holds across the market - continuous software watching, human judgment applied where the software points.
A flag is not a verdict
The most important published sentence in this whole subject area belongs to Proctorio's FAQ: Proctorio's software does not perform any type of algorithmic decision-making, such as determining whether a breach of exam integrity has occurred. All decisions regarding exam integrity are left up to the exam administrator or institution. The software's output is a mark on a timeline, not a judgment about you.
The review step is where the judgment happens, and the vendors describe it explicitly. Proctorio again: if live proctoring is in use, a proctor monitors the exam and flags certain behaviors as dictated by the exam administrator, then sends the results to the administrator for review - decisions regarding potential breaches are left to their discretion. ProctorU's published tier language is the same model in fewer words: you will be monitored and recorded during your exam, and if cheating is suspected, the instructor is notified and has video evidence of the session.
This is why a flagged moment can end in nothing. Proctorio tells test takers that subtle movements are not likely to be flagged, and that even if such actions are flagged, the institution-approved representative reviews the exam attempt and determines whether the behavior was actually a breach. The flag starts a question; a person answers it.
The flag categories vendors publish
No vendor in this source set publishes a scoring formula, but every one of them publishes the behaviors its rules target - and those behaviors are the flag categories. Proctorio names background noise directly: while you may get flagged for some background noise, Proctorio ultimately leaves it up to your exam administrator to decide whether this noise constitutes a breach of exam integrity. The same FAQ covers behavior settings, desk scans, and faces leaving the exam window.
ETS publishes a named behavior list for the GRE at home: certain behaviors will invalidate your test and can result in score cancellation - talking out loud to yourself, looking away from the screen, moving out of view of the camera, taking a break. OnVUE's session rules read as the same list from the enforcement side: do not leave the webcam view unless the exam confirms an approved break, do not speak or read aloud unless instructed, do not access your phone unless explicitly permitted by a proctor.
The Duolingo English Test's rules add the camera-frame family of categories: don't look away from the screen except when typing, take the test alone and do not interact with anyone, stay in the camera frame with ears, eyes, and mouth visible, and don't interfere with the phone camera recording. Honorlock's browser layer publishes its own blocking-and-flagging list: unauthorized websites, browser tabs, applications, and keyboard shortcuts. These are ordinary, nameable categories - nothing exotic, and nothing hidden.
Presence checks versus tracking
Proctorio's FAQ draws a line that answers one of the most common candidate fears. Proctorio does not track eye movements, the company states - but it may use facial detection to ensure test-takers are not looking away from their exam for an extended period of time. This simply detects the presence of a face interacting with the exam window. Presence is checked; gaze direction is not followed.
The same FAQ makes a parallel distinction about typing. Proctorio does not utilize keystroke fingerprinting - it does not build a profile that identifies you by how you type. What it publishes instead is that keystroke anomalies are tracked to determine whether a test-taker is consistently interacting with the exam - a consistency check that the candidate is still working, in the vendor's own general terms.
These distinctions are worth reading carefully because they are the vendors policing their own claims. A presence check answers is a face engaged with the exam. It does not answer where exactly the eyes point, who the person is behind typing patterns, or anything else. When a vendor publishes a limit like this, the limit is part of the product's public description - and this series reports the published description, nothing more.
Settings decide what gets flagged
The most useful fact for a worried candidate is that flagging is configured, not universal. Proctorio states it plainly: only the exam administrator or the institution can dictate what type of behavior they want to monitor over the course of an exam, and Proctorio will flag behavior based on the settings chosen by that administrator - settings can differ from exam to exam. Two exams on the same platform can watch for different things.
Proctorio publishes the clearest examples of what administrator configuration controls. Proctorio has no default settings, so not all test-takers will be asked to conduct a desk scan; an administrator may enable one to ensure no unauthorized resources are used. The periodic desk scan is triggered at random intervals, and institution-approved representatives evaluate whatever it captures. Recording is likewise conditional: video, audio, and screen recording happen only if the exam administrator enabled those settings.
Proctorio's product page describes the same configurability as a feature for institutions: customize monitoring with any combination of webcam, screen sharing, audio, and web traffic, with adjustable flagging, and then review sessions asynchronously or with live proctoring. For the candidate, the practical reading is that the exam's rules sheet - not the software's name - tells you what is being watched on your exam. If your institution publishes what it enables, that document is the authoritative answer.
Where the human sits: live alerts and later review
Systems differ on when the human looks, and the vendors publish the difference. Honorlock's alert model is real-time but exception-based: the software monitors and alerts a live test proctor if it detects any problems, which is why the company can tell students they won't be watched during the entire exam. ETS runs the opposite configuration for the GRE at home: the entire test session is recorded and monitored by a human proctor, who watches both you and your computer screen for the whole sitting, with suspicious movements able to invalidate the test.
The third arrangement moves the human entirely after the fact. ExamSoft's ExamMonitor is a record-and-review system, and the company's reviewer guidance describes the flow: the A.I. review runs on the recording, and a human reviewer then works through the A.I.-flagged moments - the software surfaces moments, and a person evaluates them. ProctorU's Review+ tier publishes the same shape with the instructor as the decision-maker: monitored and recorded, and if cheating is suspected, the instructor is notified and has the video evidence of the session.
None of these arrangements lets software finish the case. The live-alert systems route software output to a proctor in the session; the record-and-review systems route it to a reviewer afterward; ETS's continuous-human model barely needs the software layer described at all. In every published version, the ending is a person looking at evidence.
What the flagged test taker actually sees
The Duolingo English Test publishes the clearest candidate-facing picture of what happens after review. Every test receives the same level of human proctoring and AI scoring whether results arrive in the standard two days or the 12-hour faster-results window - paying for speed does not skip the review. The review stands behind the certification, not in front of your score only when something looks odd.
If the result cannot be certified, the outcome is visible and itemized. The test taker logs in, opens My Tests, and finds all the reasons the result was invalidated listed there, each with a learn-more article linked beside it. If a rule violation is suspected, Duolingo reserves the right not to certify the results, and the published advice is to review the test rules and requirements before retaking. A technical error in the recording - the screen recording runs the entire test - can also prevent certification.
That structure - listed reasons, linked explanations, a retake path - is the published answer to what a flag leads to on one major system. The candidate is told what was found in general terms, not merely that something went wrong. Other systems route the equivalent information through the institution instead, which is why Proctorio's FAQ repeatedly ends its answers with the administrator or institution deciding.
Why flagging exists: review is the expensive part
Honorlock publishes the number that explains the entire category: on average, you'll spend over 5 hours reviewing results for every 100 automated solution exam sessions. A proctor watching every minute of every exam does not scale, and a recording nobody reviews is just storage. Something has to tell the reviewer where to look.
That is the published job description of the flag. Proctorio's product page frames it exactly that way: focus on the moments that really matter, using adjustable flagging and analytical tools to zero in on incidents that need your judgment, and annotate results for later action when needed. The software narrows a multi-hour recording to the segments a person should evaluate; the person then brings the judgment the software is published never to make.
This is also why flag sensitivity is an administrator setting rather than a fixed constant. An institution that never wants to wade through noise can watch for fewer categories; one running a high-stakes credentialing exam can watch for more. The economics point one way for every vendor in this source set - make the human's attention expensive and well-aimed, and keep the verdict human.
What happens after a session is flagged
The published consequences attach to the human decision, not to the flag itself. On OnVUE, the session is recorded for security, quality, and training purposes, and the published penalty ladder concerns requirements and rules: failure to meet minimum requirements can bring immediate exam cancellation and forfeiture of the fee, and breaking the session rules is handled by the proctor and program. On ETS at home, violations of security procedures can result in dismissal from the test and cancellation of scores, with no refund - decided on the recorded evidence.
Before any penalty, the recording is the thing everyone looks at. ProctorU's published sentence is the pattern: if cheating is suspected, the instructor is notified and has video evidence of the session. Proctorio's FAQ adds the boundaries around that evidence: only approved representatives at your institution have access to recordings, which are stored in institution-chosen data centers, retained as the institution's agreement requires - in some cases as little as seven days - with the retention period disclosed to you during the exam pre-checks.
For the candidate, the practical sequence is stable across systems: the session runs, the software marks moments, a person reviews the recording around those moments, and only that person's decision produces an outcome - a null result, a conversation with your institution or sponsor, an invalidated score, or a cancellation with published fee consequences. A flag by itself produces none of those things.
What no vendor publishes
Set against everything the vendors do publish, the gaps are consistent. No vendor in this source set documents how a flag is computed, what threshold separates a flagged moment from an unflagged one, how any suspicion is weighted or scored, or what confidence value triggers review. Proctorio comes closest to the boundary and stops there: flagging is adjustable, settings differ, and the software makes no integrity decisions - the arithmetic, if any, is not candidate-facing documentation.
The same honesty applies to the rest of the picture. ETS publishes that it monitors reuse of devices and testing locations for unauthorized purposes, and that frequent inappropriate reuse can bring score delays, cancellation, and effects on future eligibility - without publishing how reuse is identified. Pearson VUE names assistive AI tools and leaves their workings undescribed. This page reports each claim exactly as far as the vendor's own words go and not one sentence further.
That is the correct stopping point for a general-terms explainer, and it is also the practical answer to most forum folklore. When someone claims a system flags you for some specific invisible trigger, the published record is the test: if the vendor's own pages do not say it, it is not a published fact about the system - and the published facts, categories plus human review, are enough to understand what actually happens.
If the Sitting Itself Is the Problem
AI flagging, in the published picture, is a sorting layer: software marks moments, a person decides what they mean, and the rules that decide everything are published before you book. Nothing in that picture is hidden, and nothing on this page needed to go beyond the vendors' own words.
That is the problem Exam Assist works on. If you have a proctored exam coming up and the sitting itself is the risk, send the details through the booking page: you get an honest feasibility answer, and the service fee is due only after your result posts.
For more of this series in the same general-terms format, see our guides to how OnVUE detects cheating, Proctorio's behavior monitoring, Honorlock's detection model, ExamSoft's record-and-review proctoring, Proctortrack's proctoring levels, and what to expect at a remote exam check-in.