Home › Benchmark

We tested our own detector on 1,180 human papers. It flagged 5.8% of them.

Every document in this study was published in 2018, years before ChatGPT existed. So every AI flag is a false positive. We are publishing the rate, the dataset and the method, including the parts that do not flatter us.

Documents 1,180 False positives 5.85% Published 17 August 2026 Licence CC BY 4.0

What the detector said about 1,180 human papers.

Verdict returnedDocumentsShare95% CI
Human‑Written87574.2%
Mixed / Uncertain23620.0%[17.8–22.4]
AI‑Generated695.85%[4.65–7.34]

Roughly one document in 17 was flagged as AI outright. One in four was not confidently called human at all.

These are the verdicts our product actually ships. The label comes from the detector itself, not from a threshold we picked afterwards for this page.

We went looking for the ESL bias. We did not find it.

The best-known finding in this field is Liang et al. (2023, arXiv 2304.02819), which reported that detectors misclassified more than 61% of TOEFL essays by non-native English speakers. We expected to reproduce a version of that. We did not.

GroupDocumentsFlagged as AI95% CIMedian score
Second‑language English7594.6%[3.3–6.3]28.0
Native English4218.1%[5.8–11.1]27.6

The difference is -3.5 percentage points, 95% CI [-9.9 to 2.1]. That interval contains zero, so the honest statement is that we found no significant difference in either direction. Native-English writing was flagged more often, not less, but we are not claiming that as a finding either: the interval does not support it.

Why this probably differs from the TOEFL result

Liang and colleagues tested unedited student essays. Every document here is a journal abstract that has been through peer review and copy-editing. The most likely reading is that the bias is a property of unedited second-language writing, and that editing largely removes it. That is a testable claim and we would like someone to test it.

It also means this study says nothing about student work, which is the case people actually worry about. We are not generalising to it.

The same data at every threshold.

Picking a threshold after seeing results is how a study gets flattering. Here is the false-positive rate at four score cut-offs, so you can apply your own.

Score cut-offSecond-language FPRNative FPR
≥ 505.5% [4.1–7.4]11.6% [8.9–15.1]
≥ 603.4% [2.3–5.0]7.8% [5.6–10.8]
≥ 700.8% [0.4–1.7]3.3% [2.0–5.5]
≥ 800.1% [0.0–0.7]0.0% [0.0–0.9]

By country of first-author affiliation

Countries with at least 30 documents, ordered by flag rate.

CountryGroupnFlagged as AI
United Statesnative English10613.2% [8.0–21.0]
Irelandnative English819.9% [5.1–18.3]
Australianative English519.8% [4.3–21.0]
South Koreasecond-language1017.9% [4.1–14.9]
Iransecond-language1456.9% [3.8–12.2]
New Zealandnative English814.9% [1.9–12.0]
Japansecond-language854.7% [1.8–11.5]
Brazilsecond-language1534.6% [2.2–9.1]
Spainsecond-language803.8% [1.3–10.5]
United Kingdomnative English1022.9% [1.0–8.3]
Turkeysecond-language822.4% [0.7–8.5]
Chinasecond-language422.4% [0.4–12.3]
Italysecond-language710.0% [0.0–5.1]

Method, in enough detail to rebuild it.

Where the text came from

PubMed Central's open-access subset, restricted to articles licensed CC BY or CC0 so that a commercial publisher may lawfully use and redistribute them. We deliberately did not use the standard learner corpora: ICNALE is CC BY-NC-ND, PERSUADE 2.0 and ELLIPSE are NonCommercial, and W&I+LOCNESS is not redistributable. Those licences bar commercial use, which is why you rarely see a vendor publish this measurement.

Why every flag is a false positive

Every article was published in 2018. The GPT-3 API opened in mid-2020 and ChatGPT launched in November 2022. Authorship therefore pre-dates the models, and it is timestamped by a third party rather than asserted by us. There is no need to trust anyone's claim about who wrote these documents.

How the two groups were assigned

By the first author's institutional country. Second-language group: China, Japan, South Korea, Brazil, Spain, Italy, Turkey, Iran. Native group: United States, United Kingdom, Australia, New Zealand, Ireland. Any paper carrying an affiliation from the opposing group was excluded, so a US co-author cannot place a paper in the second-language arm.

Text preparation

Abstracts only, 150–400 words. Journal structural labels ("Abstract", "Background", "Methods:") were stripped identically from both arms. This mattered: before cleaning, a leading label opened 15.7% of native documents against 4.3% of second-language ones, which would have confounded the comparison with typesetting. Non-Latin-script abstracts and results-table dumps were dropped. Median length after cleaning was 225 words for the second-language arm and 236 for the native arm, so length cannot explain the result.

Scoring

Each document was scored once through our production detection path in full mode. Determinism was verified first: the same text scored three times returned identical values, so repeat measurement was unnecessary. Documents were scored in a shuffled order with a fixed seed so both arms filled together. 1,180 of 1,180 documents scored successfully with one retry in total.

What this study does not show.

  • Affiliation country is a proxy for first language, not a measurement of it. An author in Wuhan may be a native English speaker; an author in Boston may not be. This is the weakest link in the design and we are not hiding it.
  • Copy-editing biases the gap downward. Journal abstracts are edited. That moves second-language prose toward native norms, so if a bias exists in raw writing, this design would understate it.
  • Academic abstracts only. Nothing here generalises to student essays, blog posts, emails or social media. The register is narrow and formal.
  • One year, one detector version. 2018 publications, scored on our detector as it stood in August 2026. Detector behaviour changes; this is a snapshot.
  • English only. Our detector is English-only and this study concerns English-surface text.
  • We did not test competitors. Doing so would require using their products in ways their terms forbid, so we have no comparative numbers and we are not implying any.

The data, so you can check us.

Not on request. Downloadable.

  • Per-document results (CSV, 212 KB) — one row per document with its score, verdict, word count, affiliation country, licence and a SHA-256 of the exact text scored. Every number on this page can be recomputed from this file.
  • Manifest (JSON) — counts, filters, file hashes and the limitations list in machine-readable form.

Each row carries its PMC identifier and a hash of the cleaned text, so you can re-fetch the source from PubMed Central and confirm byte-for-byte that you are scoring what we scored.

Citation

TextSight (2026). False-positive rate of the TextSight AI detector on human-written academic English. 1,180 documents. Retrieved from https://www.textsight.ai/benchmark/

Corrections

If you find an error, tell us and we will publish the correction with a date rather than quietly editing the number. Contact us.

Related: how we test the detector · why detectors produce false positives · what detection cannot do.