AI detection is full of confident percentages with no corpus behind them. We publish the dataset, the confidence intervals, the results that did not go our way, and a plain list of what we have not measured. Everything here is downloadable and reusable under CC BY 4.0.
These are commitments about how we publish, not claims about how good the product is.
Every document was published in 2018, years before ChatGPT, so every AI flag is a false positive by construction. Our detector flagged 5.85% of them. We also went looking for the well-known bias against second-language English writers and did not find it.
1,180 documents5.85% false positivesCC BY 4.0 Study 02 · 23 Aug 2026“Submit more text and the detector will get it right” is standard advice, including to people who have been wrongly accused. Between 150 and 400 words we found no relationship at all: 5.79% in the shorter half, 5.90% in the longer half, trend test p = 0.777.
5 length bandsp = 0.777 no trendCC BY 4.0One dataset currently backs both studies. It is the per-document output of the false-positive run: one row per paper, with word count, first-author affiliation group, the score the detector returned and the verdict it shipped.
Licensed CC BY 4.0. You may republish, redistribute and build on it commercially, including to argue against our conclusions, provided you credit the source.
This list exists so nobody has to guess what our silence covers. These are real gaps, stated plainly.
The data is open and the method is written to be rebuilt. If you are running a study in this area, three things may be useful:
Written by Dipak Bhosale, who runs the detection models these studies test.