If you went looking for OpenAI's GPT-2 Output Detector and got an error, that is the current state of it rather than a problem at your end. Checked on 29 September 2026: the hosted demo at openai-openai-detector.hf.space returned HTTP 503 on three consecutive attempts, and the Hugging Face Space itself reports a runtime stage of PAUSED with no hardware allocated. A paused Space does not wake up when you visit it. Somebody has to un-pause it.
The model is not gone, though, and that distinction matters. The weights at openai-community/roberta-base-openai-detector still resolve, so the detector is downloadable and self-hostable. Whether that is worth doing is the real question this page answers, because the published evidence on this tool contains two numbers that look like they contradict each other: it placed third out of fourteen tools in the largest academic comparison, and it caught 8.15% of unmodified DeepSeek-V3 output. Both are correct. Below is why.
Reported as observed HTTP responses and the Space's own published runtime state, because both are checkable and neither is a guess about intent.
| Address | Result |
|---|---|
| openai-openai-detector.hf.space (the demo) | 503, three attempts in a row |
| The Space's published runtime stage | PAUSED, allocated hardware null |
| huggingface.co/spaces/openai/openai-detector | 200, the Space page loads |
| openai-community/roberta-base-openai-detector | 200, the model weights resolve |
This is worth being precise about, because most write-ups on this tool guess. Hugging Face Spaces on free hardware sleep when idle and wake automatically on the next request, which produces a slow first load rather than an error. A paused Space is different: pausing is a deliberate action by the owner, the Space holds no hardware, and requests keep failing until someone restarts it. The runtime state published alongside the Space says PAUSED and reports current hardware as null, which is consistent with the repeated 503s rather than with a cold start.
We are not going to tell you why it was paused. No announcement was found, and a motive would be invention. The observable facts are the paused stage, the 503s, and the fact that the weights are still published.
Sources: HTTP responses observed directly on 29 September 2026; runtime stage read from the Space's own published metadata on huggingface.co. A paused Space can be restarted by its owner at any time, so this status may change after the date above.
It is older than most of the conversation around AI detection, and knowing what it was built for explains everything about how it behaves.
OpenAI released GPT-2 in stages during 2019, holding back the largest model at first over misuse concerns. Alongside that staged release it published a detector: a RoBERTa classifier fine-tuned to tell GPT-2 output apart from human text. The point was to study the detectability of its own model and give researchers something to measure with. It was never a commercial detection product, it had no threshold guidance for disciplinary use, and it predates the entire market of AI checkers that now compete on accuracy claims.
That framing matters because the tool is now used far outside its design. It is a 2019 research artifact built to recognise one specific model family, being pointed at 2026 output from models that did not exist when it was trained.
A classifier like this learns the statistical fingerprint of the generator it was trained against. GPT-2's fingerprint was comparatively coarse: more repetition, flatter structure, more obvious degeneration over long passages. Models since then produce text that is markedly more varied and more human in exactly the properties the classifier was keyed to. From the classifier's point of view, modern output looks less like the AI it knows and more like a person.
The result is not random error. It is a systematic bias toward calling things human, which is the most dangerous failure shape to be unaware of, because the tool looks reassuring rather than broken. You get a confident "human" verdict and no indication that the model has never seen anything like your input.
These two published figures are the reason this tool still has a reputation, and the reason that reputation is misleading. Reading them together is the useful part.
| Study | What it found for this tool |
|---|---|
| Weber-Wulff et al. (2023) 14 tools, 756 tests | Ranked 3rd, behind Turnitin and Compilatio |
| Alshammari and Rao (2025) 6 detectors, 294 samples | 8.15% on DeepSeek-V3 output, 93.73% on human text, 3.65% on paraphrased |
They measured different things. The two studies used different corpora, different generators and different metrics, so a ranking in one is not a rate in the other. The 2025 figures are DeepSeek-specific, and no published measurement puts this detector against current ChatGPT output at all.
More importantly, look at which numbers are high. In the 2025 study the GPT-2 Output Detector scored 93.73% on human text and 8.15% on machine text. A classifier that says "human" to almost everything will score very well on the human portion of any test set and very badly on the machine portion. If a comparison weights overall correctness across a set containing a lot of human writing, a tool biased toward "human" can place respectably while being close to useless at the job you actually want it for.
That is consistent with the Weber-Wulff paper's own conclusion, which was not a recommendation of the top finishers. Its finding was that the tools tested "are neither accurate nor reliable" and that the field's dominant bias is toward classifying output as human-written. A high placement in that study describes a tool that is less bad than the rest of a poor field, on that paper's own account.
8.15%. If you are trying to find out whether a passage was generated by a current model, a detector that catches roughly one case in twelve is not a check. And the paraphrased figure, 3.65%, is lower still: anything rewritten passes almost unconditionally.
Sources: Weber-Wulff et al., International Journal for Educational Integrity 2023, arXiv:2306.15666 (14 tools, 756 tests); Alshammari and Rao, arXiv:2507.17944 (6 detectors, 294 samples, including pre-2021 human text). The two studies are not directly comparable; each figure is reported against its own study.
The weights are published, so this is a real option. It is worth being clear about what the option is and is not.
The model repository still resolves, so the RoBERTa classifier can be downloaded and run locally or in your own Space. For a legitimate set of purposes that is genuinely useful:
OpenAI has now retired two public text detectors. The GPT-2 Output Detector's demo is paused, and the later AI Text Classifier was withdrawn on 20 July 2023 with OpenAI's own notice attributing that to "its low rate of accuracy." That classifier had published figures of 26% true positives against 9% false positives. OpenAI still describes better provenance techniques for text as something it is researching rather than something it ships. The full account is at AI Text Classifier alternative, and the related question of whether ChatGPT marks its own output is at ChatGPT watermark detector, where the short answer is that it does not.
A checklist first, then our own numbers held to the same standard, limits included.
We publish accuracy as bands rather than a headline figure: 88% to 92% on long-form English of 300 words or more across 15 model families, and 70% to 78% under about 100 words. Our false-positive rate is 5.85% across 1,180 academic papers, with per-document results downloadable at our benchmark, so you can find a document we got wrong. Output is sentence-level. The free tier is 3 checks a day with no account.
Measurement detail is at our accuracy methodology, and where we fail is listed at AI detection limitations.
Short answers, dated where the answer can change.
The hosted demo is not running. Checked on 29 September 2026, openai-openai-detector.hf.space returned HTTP 503 on three consecutive attempts, and the Hugging Face Space reports a runtime stage of PAUSED with no hardware allocated. The model weights at openai-community/roberta-base-openai-detector still resolve, so it can be downloaded and self-hosted. A paused Space can be restarted by its owner, so the demo status may change after that date.
Because it is paused rather than sleeping, and those are different states. A sleeping Space wakes automatically on the next request, which shows up as a slow first load. A paused Space holds no hardware and keeps failing until the owner restarts it. The published runtime metadata for this Space says PAUSED and reports current hardware as null, which matches the repeated 503s rather than a cold start.
Barely. Alshammari and Rao (arXiv:2507.17944) measured it at 8.15% on DeepSeek-V3 output and 3.65% on paraphrased text, while scoring 93.73% on human text. That study tested DeepSeek rather than ChatGPT, so treat it as the direction of the failure rather than a per-model figure. It was fine-tuned in 2019 to recognise GPT-2 specifically, and current models do not share that fingerprint, so it defaults to calling things human. That is a systematic bias toward false negatives rather than random error, which makes it look reassuring instead of broken.
It did, in Weber-Wulff et al. (2023), behind Turnitin and Compilatio. Both figures are true because they measured different corpora, generators and metrics. Note which numbers are high: a classifier that says "human" to nearly everything scores very well on the human portion of a test set, so it can place respectably while being poor at finding machine text. That paper's own conclusion was that the tools tested "are neither accurate nor reliable" and that the field's dominant bias is toward classifying output as human-written.
For narrow purposes, yes. It is the right choice for reproducing published results that used this exact detector, for teaching how classifier-based detection works, for detecting genuine GPT-2 output, and for offline processing where nothing may leave your machine. It is the wrong choice for judging current model output, and self-hosting cannot change what the model was trained on. It also ships no threshold guidance, because it was never designed to support a decision about a person.
Yes. Its later AI Text Classifier was withdrawn on 20 July 2023, with OpenAI's own notice attributing the withdrawal to "its low rate of accuracy"; its published figures were 26% true positives against 9% false positives. OpenAI continues to describe more effective provenance techniques for text as something it is researching rather than something it ships, and it has never released a text watermark, though it built one.
Something fitted to models currently in use, that publishes a false-positive rate against a named corpus, returns sentence-level output rather than one number, and states its own weak cases. Ours publishes bands of 88% to 92% on long-form English of 300 words or more and 70% to 78% under about 100 words, plus a 5.85% false-positive rate over 1,180 academic papers with the per-document data downloadable. It is English only and it carries the same bias against second-language English the rest of the field does.
The other OpenAI detector, withdrawn July 2023 with its own low-accuracy notice.
Read the history →The watermark OpenAI built, never shipped, and why there is nothing to detect.
Read the answer →Another detector gone from self-service, reported as a redirect rather than a motive.
See the status →1,180 academic papers, 5.85%, per-document data you can download and audit.
Check our numbers →Where statistical detection fails, listed by the people selling it to you.
Read the limits →How the bands were measured, on what data, and why there is no single figure.
Read the method →Sentence-level highlights, accuracy published as bands rather than a headline number, and a false-positive benchmark you can download and audit. 3 checks a day, no account, no card.