HomeResources › AI Detector Accuracy Explained

AI detector accuracy, explained honestly.

Every detector publishes an accuracy number. Almost none of them mean what readers think they mean. This guide walks through what TPR and FPR actually measure, why ESL writing gets over-flagged, what the peer-reviewed literature says, and which tools hold up on identical passages. No marketing language, no hidden bias, no claim that any one tool is always right.

Try TextSight free
1,180 papers in our own benchmark Per-document dataset downloadable Peer-reviewed sources cited Last verified
The short answer

What "accurate" really means here.

Three things to keep in mind before you read any vendor's claim about detection accuracy.

One. A vendor accuracy number is almost always a true positive rate measured on a vendor-chosen test set. It says how often the tool catches AI on writing similar to that set. It does not say how often it wrongly flags a real student.

Two. The number that determines whether you can trust a verdict is the false positive rate, broken out by writer type. A 1% FPR on native English can sit next to a 22% FPR on ESL writing inside the same tool. Vendors rarely publish that split. Independent researchers do.

Three. Every accuracy number degrades the moment a paraphraser, a round of human editing, or an unfamiliar topic enters the picture. Treat any score as a probability, not a fact.

Bottom line. The best detectors in 2026 sit around 90 to 97 percent TPR on raw AI and 2 to 6 percent FPR on careful native English. On ESL writing the published spread is far wider: Liang et al. measured a 61.3% average across seven detectors on TOEFL essays, and our own rate on second-language writers in our benchmark set was 4.6%, against 8.1% for native writers, a gap whose confidence interval crosses zero. No single verdict should be treated as proof.
The two numbers that matter

TPR and FPR, in plain English.

Every detector pitch hides one of these two numbers. Reading them together is the only honest way to evaluate a tool.

True positive rate (TPR), or "did it catch the AI"

TPR is the share of AI-generated passages the tool correctly flags. A 92% TPR means the detector catches 92 out of every 100 AI samples. Vendors love this number because it is easy to push up. It says nothing about cost. A tool with 99% TPR and a 30% FPR is worse than one with 91% TPR and a 3% FPR for any real classroom.

False positive rate (FPR), or "how often it accuses the innocent"

FPR is the share of human-written passages wrongly flagged as AI. This is the number that determines whether you can trust a verdict. On a class of 30 essays, a 5% FPR means roughly 1.5 students wrongly accused. A 22% FPR, which Stanford measured on TOEFL essays in 2023, means closer to 7. The cost of a false positive falls on the writer.

The trade-off curve nobody shows you

Every detector has one knob: the threshold. Lower it and TPR rises but FPR rises too. Raise it and FPR drops but TPR drops with it. The vendor's headline number is whichever point on the curve makes the marketing read best. To compare honestly, fix the threshold or compare full curves.

Recall, precision, and the marketing fog

ML literature calls TPR "recall" and pairs it with "precision" (of everything flagged, how much was AI). Marketing pages translate these into "accuracy" or "detection rate" without disclosing the test set or threshold. If a page says "99% accurate" without splitting TPR and FPR, the writer either does not know the difference or is hoping you do not.

The bias problem

Why detectors flag ESL writers more often.

The single most documented failure mode in the academic literature. Worth understanding before trusting any verdict that involves a non-native English writer.

The Stanford finding

In July 2023, Liang and colleagues at Stanford published a paper in Patterns (Cell Press) measuring detector accuracy on TOEFL essays by non-native English speakers. More than half were misclassified as AI-generated by mainstream detectors, with one configuration reaching 61% FPR. On essays from US-born eighth-graders, the same detectors held under 5%. The model was reading the structural footprint of formally-taught English as machine-generated.

Why this happens at the signal level

Classical detectors score perplexity and burstiness. Second-language academic writing uses a constrained vocabulary, follows taught templates, and produces uniform sentence lengths. All three reduce perplexity and burstiness. The signal the detector reads as "machine" overlaps the signal of a careful non-native writer. Not a bug in any one tool: a property of the underlying method.

What's been done about it

Most major vendors have re-tuned since the Stanford paper. GPTZero shipped a 2024 ESL update. Turnitin recalibrated threshold defaults. Originality.ai added a language-aware second pass. Gains are uneven, and we are not aware of a peer-reviewed retest since Weber-Wulff and Liang that measures the field as a whole. TextSight is calibrated to keep its false-positive rate low on the same writing.

The practical takeaway

If you write in formally-taught English (Indian-curriculum, Filipino academic, Chinese university register), expect more false flags than the vendor's headline number predicts. Pre-scan before submission. Re-scan any flag on a second independent tool. If you teach or grade ESL essays, treat any single verdict as a starting point for a conversation, not evidence.

Pattern recognition

Why careful human writing gets flagged.

Low perplexity, low burstiness, formulaic transitions, technical register, and short length are the five surface properties that push human writing toward an AI verdict. They are statistical features of the text, not stylistic crimes.

We cover each of the five in detail, with what to do about them, on the page built for that question: AI detector false positives, and who gets caught by them.

The short version for this page is that all five describe the same underlying situation. A detector estimates how surprising your next word is, given the words before it. Writing that is careful, edited, formally taught, or technical is less surprising by construction, and a model reads low surprise as machine generation. That is why the error lands hardest on second-language writers and STEM students, and why a low score is not a compliment about your prose.

The field, honestly

What the peer-reviewed evidence actually establishes.

We have not independently tested other vendors' detectors under controlled conditions, so we publish no false-positive numbers for them. What follows is what the published literature measured, and what we measured on our own detector with the dataset attached.

Peer-reviewed findings and our own published measurement · last checked 2026-09-09
Finding Measured rate Source
False positives on TOEFL essays by non-native English writers, averaged across seven GPT detectors (Originality.AI, Quil.org, Sapling, OpenAI, Crossplag, GPTZero, ZeroGPT)61.3%Liang et al., Patterns (Cell Press), 2023
The same seven detectors on essays by native English-speaking US eighth-gradersnear zeroLiang et al., 2023
Cross-detector audit of 14 tools, including Turnitin, across human, machine, and machine-paraphrased text. No tool classified machine-paraphrased text reliablysee paperWeber-Wulff et al., Int. J. Educational Integrity, 2023
TextSight, on 1,180 pre-ChatGPT human-written papers5.85% [4.65–7.34]Our own benchmark, per-document dataset downloadable
TextSight, second-language against native writers within that set4.6% vs 8.1%The confidence interval crosses zero, so we claim no significant difference rather than an advantage

Vendors publish their own false-positive claims in their documentation. Those are their claims, not measurements we have verified, so we do not reproduce them here as though they were ours. A vendor appears above only where a peer-reviewed study named it. Our own figure and its per-document dataset are at the benchmark, and how we test is at our methodology page.

The protocol

If a detector flagged your human writing.

A five-step protocol used by students, teachers, and editors to dispute a wrong AI flag. Practical, not adversarial.

Step 1: Preserve drafts before you touch anything

The instinct after a flag is to rewrite. Resist it. A panic rewrite destroys version history, your strongest evidence. Capture Google Docs revision history, Word AutoRecover, Notion page history, or browser autosave. Edit timestamps showing 40 minutes of incremental revision are the closest thing to proof.

Step 2: Re-scan on a second independent detector

One verdict is a probability. Two agreeing is a stronger signal. Run the passage through a detector using a different signal family: if the first was GPTZero (perplexity), run TextSight (sentence rhythm). Disagreement is itself evidence the verdict is not reliable.

Step 3: Request the methodology and threshold

Any reviewer using a detector verdict against you owes three things: the published methodology, the threshold (50%, 60%, 80% confidence floor), and the per-sentence breakdown. Most academic integrity policies require this disclosure on request. Refusal is grounds for escalation.

Step 4: Bring the per-sentence breakdown to the appeal

Modern detectors show which sentences scored high and why. If the flagged sentences carry formulaic transitions, learned templates, or technical register, that pattern alone explains the flag and is worth naming. "This sentence flagged because it has low burstiness, a structural property of careful academic prose" lands differently than a generic denial.

A note on tone. Appeals work best when they treat the detector as a flawed instrument rather than the reviewer as a bad actor. Most teachers want the tool to work. Show them why the verdict is not safe in this case.
FAQ

AI detector accuracy, frequently asked.

How accurate are AI detectors in 2026?
It depends. On raw GPT-4 or Claude the top detectors land between 88% and 97% TPR. On paraphrased AI, accuracy often drops into 40 to 60 percent. On human writing, field FPR ranges between 1% and 22% depending on first language, register, and threshold. No single tool is reliable enough to be treated as proof.
What do TPR and FPR mean for an AI detector?
TPR is the share of AI passages a detector correctly flags. FPR is the share of human passages it wrongly flags. A 99% TPR can sit next to a 30% FPR. The cost of a wrong flag falls on the writer, so FPR is the number to read first.
Why do AI detectors flag ESL writing so often?
Second-language academic writing has lower perplexity and burstiness than native English, the exact signals classical detectors read as machine. Liang et al. at Stanford (2023) measured 61% FPR on TOEFL essays across multiple detectors.
Can a school punish a student based on a detector score alone?
Reputable academic integrity guidance, including from Turnitin and GPTZero, says no detector output should be the sole basis for disciplinary action. Ask for the methodology, the threshold, and the per-sentence breakdown.
Do AI detectors work on paraphrased or humanized text?
Most detectors lose significant recall on paraphrased AI. On lightly humanized passages, TextSight holds up better than detectors that lean on the older surface tells. Perplexity scoring degrades fastest because paraphrasers add the exact variance the detector reads.
What should I do if a detector wrongly flags my writing?
Do not panic-rewrite. Preserve drafts and version history. Re-scan on a second independent detector with a different signal family. Request the methodology and threshold. Bring the per-sentence breakdown to any appeal.
Related

More on detector accuracy and methodology.

Further reading

Run a passage through TextSight. Read the per-sentence evidence.

Free tier: 3 scans a day, 5,000 characters per scan, no card, no email, no signup. The fastest way to test the accuracy claims on this page against your own writing.

Start free, no card Read the methodology
Sentence-level highlights · ESL-aware false-positive tuning · Peer-reviewed sources cited · No signup required for the free tier

AI detection, more places & platforms

How TextSight works for other regions and setups.

AI Detector False Positives Explained AI Detector for Argentina, Argentine English AI Detector for Australia: Go8 & Turnitin Ready AI Detector for Bangladesh, Bangladeshi English AI Detector for Belgium: KU Leuven, Brussels EU AI Detector for Brazil, Brazilian English

Our own accuracy figures, with the data.

Rather than quote a headline accuracy percentage, we publish the measurements themselves — including the ones that do not flatter us.

  • False positives on 1,180 human papers — every document published in 2018, years before ChatGPT, so every AI flag is a false positive by construction. Ours flagged 5.85% [4.65–7.34]. We also looked for the well-known bias against second-language English writers and did not find a significant one.
  • Does length protect you? — between 150 and 392 words, no. 5.79% in the shorter half against 5.90% in the longer half, trend test p = 0.777.

Both come with the full per-document dataset under CC BY 4.0, so the numbers can be checked rather than taken on trust. See all research.