Every detector publishes an accuracy number. Almost none of them mean what readers think they mean. This guide walks through what TPR and FPR actually measure, why ESL writing gets over-flagged, what the peer-reviewed literature says, and which tools hold up on identical passages. No marketing language, no hidden bias, no claim that any one tool is always right.
Three things to keep in mind before you read any vendor's claim about detection accuracy.
One. A vendor accuracy number is almost always a true positive rate measured on a vendor-chosen test set. It says how often the tool catches AI on writing similar to that set. It does not say how often it wrongly flags a real student.
Two. The number that determines whether you can trust a verdict is the false positive rate, broken out by writer type. A 1% FPR on native English can sit next to a 22% FPR on ESL writing inside the same tool. Vendors rarely publish that split. Independent researchers do.
Three. Every accuracy number degrades the moment a paraphraser, a round of human editing, or an unfamiliar topic enters the picture. Treat any score as a probability, not a fact.
Every detector pitch hides one of these two numbers. Reading them together is the only honest way to evaluate a tool.
TPR is the share of AI-generated passages the tool correctly flags. A 92% TPR means the detector catches 92 out of every 100 AI samples. Vendors love this number because it is easy to push up. It says nothing about cost. A tool with 99% TPR and a 30% FPR is worse than one with 91% TPR and a 3% FPR for any real classroom.
FPR is the share of human-written passages wrongly flagged as AI. This is the number that determines whether you can trust a verdict. On a class of 30 essays, a 5% FPR means roughly 1.5 students wrongly accused. A 22% FPR, which Stanford measured on TOEFL essays in 2023, means closer to 7. The cost of a false positive falls on the writer.
Every detector has one knob: the threshold. Lower it and TPR rises but FPR rises too. Raise it and FPR drops but TPR drops with it. The vendor's headline number is whichever point on the curve makes the marketing read best. To compare honestly, fix the threshold or compare full curves.
ML literature calls TPR "recall" and pairs it with "precision" (of everything flagged, how much was AI). Marketing pages translate these into "accuracy" or "detection rate" without disclosing the test set or threshold. If a page says "99% accurate" without splitting TPR and FPR, the writer either does not know the difference or is hoping you do not.
The single most documented failure mode in the academic literature. Worth understanding before trusting any verdict that involves a non-native English writer.
In July 2023, Liang and colleagues at Stanford published a paper in Patterns (Cell Press) measuring detector accuracy on TOEFL essays by non-native English speakers. More than half were misclassified as AI-generated by mainstream detectors, with one configuration reaching 61% FPR. On essays from US-born eighth-graders, the same detectors held under 5%. The model was reading the structural footprint of formally-taught English as machine-generated.
Classical detectors score perplexity and burstiness. Second-language academic writing uses a constrained vocabulary, follows taught templates, and produces uniform sentence lengths. All three reduce perplexity and burstiness. The signal the detector reads as "machine" overlaps the signal of a careful non-native writer. Not a bug in any one tool: a property of the underlying method.
Most major vendors have re-tuned since the Stanford paper. GPTZero shipped a 2024 ESL update. Turnitin recalibrated threshold defaults. Originality.ai added a language-aware second pass. Gains are uneven, and we are not aware of a peer-reviewed retest since Weber-Wulff and Liang that measures the field as a whole. TextSight is calibrated to keep its false-positive rate low on the same writing.
If you write in formally-taught English (Indian-curriculum, Filipino academic, Chinese university register), expect more false flags than the vendor's headline number predicts. Pre-scan before submission. Re-scan any flag on a second independent tool. If you teach or grade ESL essays, treat any single verdict as a starting point for a conversation, not evidence.
Low perplexity, low burstiness, formulaic transitions, technical register, and short length are the five surface properties that push human writing toward an AI verdict. They are statistical features of the text, not stylistic crimes.
We cover each of the five in detail, with what to do about them, on the page built for that question: AI detector false positives, and who gets caught by them.
The short version for this page is that all five describe the same underlying situation. A detector estimates how surprising your next word is, given the words before it. Writing that is careful, edited, formally taught, or technical is less surprising by construction, and a model reads low surprise as machine generation. That is why the error lands hardest on second-language writers and STEM students, and why a low score is not a compliment about your prose.
We have not independently tested other vendors' detectors under controlled conditions, so we publish no false-positive numbers for them. What follows is what the published literature measured, and what we measured on our own detector with the dataset attached.
| Finding | Measured rate | Source |
|---|---|---|
| False positives on TOEFL essays by non-native English writers, averaged across seven GPT detectors (Originality.AI, Quil.org, Sapling, OpenAI, Crossplag, GPTZero, ZeroGPT) | 61.3% | Liang et al., Patterns (Cell Press), 2023 |
| The same seven detectors on essays by native English-speaking US eighth-graders | near zero | Liang et al., 2023 |
| Cross-detector audit of 14 tools, including Turnitin, across human, machine, and machine-paraphrased text. No tool classified machine-paraphrased text reliably | see paper | Weber-Wulff et al., Int. J. Educational Integrity, 2023 |
| TextSight, on 1,180 pre-ChatGPT human-written papers | 5.85% [4.65–7.34] | Our own benchmark, per-document dataset downloadable |
| TextSight, second-language against native writers within that set | 4.6% vs 8.1% | The confidence interval crosses zero, so we claim no significant difference rather than an advantage |
Vendors publish their own false-positive claims in their documentation. Those are their claims, not measurements we have verified, so we do not reproduce them here as though they were ours. A vendor appears above only where a peer-reviewed study named it. Our own figure and its per-document dataset are at the benchmark, and how we test is at our methodology page.
A five-step protocol used by students, teachers, and editors to dispute a wrong AI flag. Practical, not adversarial.
The instinct after a flag is to rewrite. Resist it. A panic rewrite destroys version history, your strongest evidence. Capture Google Docs revision history, Word AutoRecover, Notion page history, or browser autosave. Edit timestamps showing 40 minutes of incremental revision are the closest thing to proof.
One verdict is a probability. Two agreeing is a stronger signal. Run the passage through a detector using a different signal family: if the first was GPTZero (perplexity), run TextSight (sentence rhythm). Disagreement is itself evidence the verdict is not reliable.
Any reviewer using a detector verdict against you owes three things: the published methodology, the threshold (50%, 60%, 80% confidence floor), and the per-sentence breakdown. Most academic integrity policies require this disclosure on request. Refusal is grounds for escalation.
Modern detectors show which sentences scored high and why. If the flagged sentences carry formulaic transitions, learned templates, or technical register, that pattern alone explains the flag and is worth naming. "This sentence flagged because it has low burstiness, a structural property of careful academic prose" lands differently than a generic denial.
Source-cited audit of GPTZero's claims against independent academic findings.
Read the review →Five writing styles that trigger false flags, plus a protocol if you were wrongly flagged.
Read the playbook →Mechanism guide to perplexity, burstiness, and rhythm-based scoring failures.
Read the breakdown →Turnitin's 4% FPR claim against measured field results on ESL and STEM writing.
Read the review →Head-to-head detection benchmark with pricing and ESL false positives compared.
Read the compare →The benchmark corpus, threshold definitions, and reproducibility notes.
Read the methodology →A template for responding when a detector flags work you wrote, and the errors that sink a case.
Use the template →Free tier: 3 scans a day, 5,000 characters per scan, no card, no email, no signup. The fastest way to test the accuracy claims on this page against your own writing.
How TextSight works for other regions and setups.
Rather than quote a headline accuracy percentage, we publish the measurements themselves — including the ones that do not flatter us.
Both come with the full per-document dataset under CC BY 4.0, so the numbers can be checked rather than taken on trust. See all research.