Every document in this study was published in 2018, years before ChatGPT existed. So every AI flag is a false positive. We are publishing the rate, the dataset and the method, including the parts that do not flatter us.
| Verdict returned | Documents | Share | 95% CI |
|---|---|---|---|
| Human‑Written | 875 | 74.2% | — |
| Mixed / Uncertain | 236 | 20.0% | [17.8–22.4] |
| AI‑Generated | 69 | 5.85% | [4.65–7.34] |
Roughly one document in 17 was flagged as AI outright. One in four was not confidently called human at all.
These are the verdicts our product actually ships. The label comes from the detector itself, not from a threshold we picked afterwards for this page.
The best-known finding in this field is Liang et al. (2023, arXiv 2304.02819), which reported that detectors misclassified more than 61% of TOEFL essays by non-native English speakers. We expected to reproduce a version of that. We did not.
| Group | Documents | Flagged as AI | 95% CI | Median score |
|---|---|---|---|---|
| Second‑language English | 759 | 4.6% | [3.3–6.3] | 28.0 |
| Native English | 421 | 8.1% | [5.8–11.1] | 27.6 |
The difference is -3.5 percentage points, 95% CI [-9.9 to 2.1]. That interval contains zero, so the honest statement is that we found no significant difference in either direction. Native-English writing was flagged more often, not less, but we are not claiming that as a finding either: the interval does not support it.
Liang and colleagues tested unedited student essays. Every document here is a journal abstract that has been through peer review and copy-editing. The most likely reading is that the bias is a property of unedited second-language writing, and that editing largely removes it. That is a testable claim and we would like someone to test it.
It also means this study says nothing about student work, which is the case people actually worry about. We are not generalising to it.
Picking a threshold after seeing results is how a study gets flattering. Here is the false-positive rate at four score cut-offs, so you can apply your own.
| Score cut-off | Second-language FPR | Native FPR |
|---|---|---|
| ≥ 50 | 5.5% [4.1–7.4] | 11.6% [8.9–15.1] |
| ≥ 60 | 3.4% [2.3–5.0] | 7.8% [5.6–10.8] |
| ≥ 70 | 0.8% [0.4–1.7] | 3.3% [2.0–5.5] |
| ≥ 80 | 0.1% [0.0–0.7] | 0.0% [0.0–0.9] |
Countries with at least 30 documents, ordered by flag rate.
| Country | Group | n | Flagged as AI |
|---|---|---|---|
| United States | native English | 106 | 13.2% [8.0–21.0] |
| Ireland | native English | 81 | 9.9% [5.1–18.3] |
| Australia | native English | 51 | 9.8% [4.3–21.0] |
| South Korea | second-language | 101 | 7.9% [4.1–14.9] |
| Iran | second-language | 145 | 6.9% [3.8–12.2] |
| New Zealand | native English | 81 | 4.9% [1.9–12.0] |
| Japan | second-language | 85 | 4.7% [1.8–11.5] |
| Brazil | second-language | 153 | 4.6% [2.2–9.1] |
| Spain | second-language | 80 | 3.8% [1.3–10.5] |
| United Kingdom | native English | 102 | 2.9% [1.0–8.3] |
| Turkey | second-language | 82 | 2.4% [0.7–8.5] |
| China | second-language | 42 | 2.4% [0.4–12.3] |
| Italy | second-language | 71 | 0.0% [0.0–5.1] |
PubMed Central's open-access subset, restricted to articles licensed CC BY or CC0 so that a commercial publisher may lawfully use and redistribute them. We deliberately did not use the standard learner corpora: ICNALE is CC BY-NC-ND, PERSUADE 2.0 and ELLIPSE are NonCommercial, and W&I+LOCNESS is not redistributable. Those licences bar commercial use, which is why you rarely see a vendor publish this measurement.
Every article was published in 2018. The GPT-3 API opened in mid-2020 and ChatGPT launched in November 2022. Authorship therefore pre-dates the models, and it is timestamped by a third party rather than asserted by us. There is no need to trust anyone's claim about who wrote these documents.
By the first author's institutional country. Second-language group: China, Japan, South Korea, Brazil, Spain, Italy, Turkey, Iran. Native group: United States, United Kingdom, Australia, New Zealand, Ireland. Any paper carrying an affiliation from the opposing group was excluded, so a US co-author cannot place a paper in the second-language arm.
Abstracts only, 150–400 words. Journal structural labels ("Abstract", "Background", "Methods:") were stripped identically from both arms. This mattered: before cleaning, a leading label opened 15.7% of native documents against 4.3% of second-language ones, which would have confounded the comparison with typesetting. Non-Latin-script abstracts and results-table dumps were dropped. Median length after cleaning was 225 words for the second-language arm and 236 for the native arm, so length cannot explain the result.
Each document was scored once through our production detection path in full mode. Determinism was verified first: the same text scored three times returned identical values, so repeat measurement was unnecessary. Documents were scored in a shuffled order with a fixed seed so both arms filled together. 1,180 of 1,180 documents scored successfully with one retry in total.
Not on request. Downloadable.
Each row carries its PMC identifier and a hash of the cleaned text, so you can re-fetch the source from PubMed Central and confirm byte-for-byte that you are scoring what we scored.
TextSight (2026). False-positive rate of the TextSight AI detector on human-written academic English. 1,180 documents. Retrieved from https://www.textsight.ai/benchmark/
If you find an error, tell us and we will publish the correction with a date rather than quietly editing the number. Contact us.
Related: how we test the detector · why detectors produce false positives · what detection cannot do.