Home › Resources › Is ZeroGPT Accurate?

Is ZeroGPT accurate? The honest 2026 answer.

Sometimes, and you cannot tell which times from the score alone. ZeroGPT's own homepage advertises 98.4% detection accuracy and a false positive rate under 1%. Those are strong numbers. They are also unsourced: ZeroGPT publishes no methodology paper, names no evaluation set, and cites no independent validation anywhere on its site. It calls the underlying method "DeepAnalyse™ Technology" and leaves it there.

ZeroGPT was tested independently, in the largest academic comparison of detectors to date. It was one of 14 tools in Weber-Wulff et al. (2023), and that paper's conclusion is blunt: the available tools "are neither accurate nor reliable." Below: what ZeroGPT claims, what the independent record measured, the confusion with GPTZero that sends people to the wrong tool, and what to do if ZeroGPT flagged your writing.

Run a second-opinion scan free
3 checks/day free No signup required Source-cited audit Last verified
The published claims

What ZeroGPT actually claims.

Start with the numbers the tool publishes about itself, because almost every argument about ZeroGPT is really an argument about how much those numbers are worth.

The headline numbers

ZeroGPT's homepage states 98.4% detection accuracy and a false positive rate below 1%, alongside claimed coverage of "20+ LLM families." The free text box accepts up to 15,000 characters per check, which is generous compared with most free detectors. Paid plans exist, with a 30% annual discount advertised, and a permanently free tier that needs no credit card.

Those two headline figures would put ZeroGPT among the best detectors ever built if they held up in the open. The problem is not that they are implausible. The problem is that there is nothing attached to them.

What is missing from the claim

A detector accuracy number is only meaningful with three things beside it: the evaluation set, the decision threshold, and the generator mix. ZeroGPT publishes none of them. There is no methodology paper, no named benchmark, no dataset, no per-model breakdown, and no citation to a peer-reviewed evaluation. "DeepAnalyse™ Technology" is described as "a pioneering research in the modeling of AI content detection," which is marketing language, not a method.

Compare that with GPTZero, which has published a methodology write-up, named its known weak cases (ESL writing, passages under 250 words, paraphrased text), and tells institutions in writing not to treat its verdict as decisive. You can disagree with GPTZero's numbers, but you can at least locate them. With ZeroGPT there is nothing to check.

How accurate is ZeroGPT on false positives, really?

False positive rate is the number that matters to a student or a freelancer, because it is the number that decides whether clean human writing gets flagged. It is also the easiest number to make look good: measure it on a corpus of casual, bursty, native-English prose and almost any detector will report under 1%. Measure it on formal academic writing by a second-language author and the same model can flag one in five.

So a sub-1% figure with no stated corpus is not a claim about your writing. It is a claim about someone else's, and you were probably not in that sample.

Sources: zerogpt.com homepage and pricing page, claims as displayed September 2026. No methodology paper or third-party validation is cited on either page.

The independent record

ZeroGPT accuracy: what independent testing actually found.

ZeroGPT has been named in two peer-reviewed evaluations, and neither is on its own website. The larger of the two is the single most useful thing on this page.

Weber-Wulff et al. (2023): ZeroGPT was in the sample

Published in the International Journal for Educational Integrity, this nine-author international study ran 756 tests across 14 detection tools, ZeroGPT among them, using an original document set the authors built themselves. It remains the most comprehensive independent comparison of AI text detectors in the literature.

Its conclusion, verbatim from the abstract: the available detection tools "are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques significantly worsen the performance of tools."

Read that bias direction carefully, because it is the opposite of what most people assume. Across the field, the aggregate failure mode was missing AI text, not over-flagging human text.

The numbers, by category

The study reported accuracy by document type rather than publishing a league table for every tool. These are field-wide averages across all 14 tools, which is exactly the right frame for judging an unsourced single-tool claim:

Document typeAverage accuracy across the 14 tools
Human-written text~95%
Machine-translated text~75%
Unmodified ChatGPT output~74%
AI text after manual editing~42%
AI text paraphrased with QuillBot~26%

And the miss rates

Stated the other way round: AI-generated text went undetected in roughly 20% of unmodified cases, 52% after manual revision, and 71% after a pass through QuillBot. False positives averaged about 2% across the field, rising to about 11% once machine translation was involved.

On the paper's own ranking, Turnitin scored highest of the 14 tools, followed by Compilatio and the GPT-2 Output Detector. ZeroGPT was not in that top group. The authors' overall verdict was that none of the tested tools reported AI-generated text with satisfactory accuracy.

Source: Weber-Wulff et al. (2023), Weber-Wulff, Anohina-Naumeca, et al., "Testing of detection tools for AI-generated text," International Journal for Educational Integrity, 2023 (arXiv:2306.15666). Category figures as summarised by Technische Hochschule Mittelhessen.

Liang et al., Stanford (2023): the ESL problem is not ZeroGPT-specific

The Stanford team ran seven detectors against TOEFL essays written by non-native English speakers and measured an average false positive rate of 61.3%, against a near-zero rate on essays by native English-speaking US eighth-graders. ZeroGPT was one of the seven (the full list: Originality.AI, Quil.org, Sapling, OpenAI’s classifier, Crossplag, GPTZero and ZeroGPT), so unlike most claims on this page, this one is a measurement that includes the tool itself. Second-language academic prose has lower perplexity and lower burstiness, which are the exact two properties classical detectors read as "AI-like."

This is a structural calibration problem across the whole category, ours included. No detector has solved it, and any tool advertising a sub-1% false positive rate without naming its corpus has almost certainly not measured against writing like this.

Source: Liang et al., "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023.

The name trap

ZeroGPT is not GPTZero.

Two different companies, two different products, two nearly identical names. A large share of "is ZeroGPT accurate" searches are actually about the other one.

GPTZero (gptzero.me) was built by Edward Tian in early 2023, is the tool most universities reference, publishes a methodology, and currently advertises 99% accuracy with a 10,000-character free input. ZeroGPT (zerogpt.com) is a separate commercial product with a 15,000-character free box, the 98.4% claim, and no published methodology.

This matters practically. If your professor said "the ZeroGPT result," ask which URL they used before you accept anything about the result, because the two tools disagree on the same passage often enough to change an outcome. We keep the full spec-by-spec comparison, including both companies' self-reported false-positive claims and what each does and does not disclose, on a dedicated page: GPTZero vs ZeroGPT.

Honest credit

Where ZeroGPT is genuinely useful.

A fair audit names the cases where the tool earns its traffic. There are three.

Raw, unedited long-form AI output

Paste 800 words straight out of a chatbot with no editing pass and ZeroGPT will usually call it. This is the case every detector is built for and the one where the whole field performs best, around 74% on average in the Weber-Wulff set. If you want a fast sanity check on an inbound draft that looks machine-written, it is a reasonable first look.

A free 15,000-character box with no signup

The free allowance is genuinely one of the most generous in the category, and there is no account wall before your first result. For a one-off check that is real value, and it is most of why the tool is so widely used.

As one of several signals, never as the verdict

Used the way the literature recommends, as a prompt to look closer rather than a finding, ZeroGPT is fine. The damage happens at the next step, when an unsourced percentage gets treated as evidence in a misconduct meeting.

The real failure modes

Where ZeroGPT falls down.

Four cases where the 98.4% figure does not generalise. If your writing sits in any of them, treat any score as provisional.

Second-language English in academic register

The most consequential failure case in the whole category, and the one with the strongest evidence behind it. Stanford measured a 61.3% average false positive rate on non-native TOEFL essays across seven detectors, and ZeroGPT was one of the seven tested. The paper reports the seven-tool average rather than a per-tool breakdown, so the exact ZeroGPT figure is not public, but the tool was in the sample that produced that average. ZeroGPT publishes nothing about how it handles ESL writing.

Anything that has been through a paraphraser

This is the field-wide collapse the Weber-Wulff data quantifies: accuracy fell to about 26% on QuillBot-paraphrased AI text, with 71% of it going undetected. Content obfuscation "significantly worsen[s] the performance of tools," in the authors' words. A detector that cannot survive one paraphrasing pass cannot support a disciplinary decision.

Short passages

Perplexity and burstiness statistics need text to stabilise. Under roughly 250 words, scores get noisy enough that the same paragraph can return materially different results on consecutive scans. GPTZero documents a 250-word floor for this reason; ZeroGPT publishes no minimum, which does not mean it does not have one.

Technical, formulaic and translated prose

Lab reports, clinical write-ups, mathematical exposition and patent text all share low burstiness because the register demands repeatable phrasing. Human authors in those genres get flagged routinely across every detector. Machine-translated text is its own trap: the Weber-Wulff set showed false positives roughly five times higher, about 11%, once translation was in the pipeline.

What to do about it

Flagged by ZeroGPT? A four-step protocol.

An unsourced percentage is not evidence. This is the sequence that has actually worked for people who wrote their own work and got flagged anyway.

Step 1: ask which tool, and ask for the threshold

Establish whether it was ZeroGPT or GPTZero, and ask in writing what score the institution treats as actionable. If no threshold is documented, that is worth knowing early, because it means the decision is being made on vibes rather than a published policy.

Step 2: re-scan on an independent detector

Two detectors built on different signals disagreeing on the same passage is meaningful and easy to demonstrate. Run the identical text through a second tool and keep both screenshots side by side, with timestamps. Our free scan takes 1,000 words with no account, which is enough for most single-assignment disputes.

Step 3: produce your process, not your prose

Version history is the strongest evidence a writer has, and it is the one thing no detector can manufacture. Google Docs revision history, Word autosaves, git commits, notes, outlines, browser history on your sources. A document that visibly grew over eleven sittings did not arrive from a chatbot.

Step 4: put the literature in front of them

Most people making these decisions have never read a detector evaluation. Cite the peer-reviewed record rather than arguing about your own score: the Weber-Wulff study concluding no tool is satisfactorily accurate, and the Stanford ESL finding if English is your second language. We keep an appeal-letter template that does this properly.

Our own position

What we publish, and what we refuse to.

We sell a detector, so treat this section with the scepticism it deserves. Here is what we will and will not put a number on.

We do not publish a single headline accuracy percentage, deliberately. One number across every length, genre and register is misleading, and it is precisely the move this page has spent 2,000 words criticising. What we publish instead are bands: on long-form English of roughly 300 words or more, against 15 current commercial model families, our internal benchmark sits at roughly 88 to 92% accuracy. Under about 100 words it drops to roughly 70 to 78%. Those are our own internal results, not an independent evaluation, and we say so on the page they live on.

On false positives, we published the measurement with the dataset attached: 1,180 academic papers, 5.85% measured false positive rate, downloadable per document so you can check our arithmetic. That number is worse than ZeroGPT's advertised sub-1%. It is also the only one of the two you can audit.

The detector is English-only. It carries the same documented second-language false-positive risk as every tool in this category, which is why we include non-native academic prose in training and surface a low-confidence flag on borderline samples rather than rounding them to a clean verdict. We do not have an independent third-party benchmark, we want one, and we will not claim one until it exists.

Sources: our accuracy methodology page and our published false-positive benchmark. Both link the underlying dataset.

Questions

Is ZeroGPT accurate, frequently asked.

Is ZeroGPT accurate enough to be used as evidence of cheating?

No. No detector verdict should be used as standalone evidence, and ZeroGPT is a weaker case than most because it publishes no methodology, no evaluation set and no threshold guidance you could examine in an appeal. The peer-reviewed comparison that tested it across 756 tests concluded that the available tools are "neither accurate nor reliable." Treat any score as a reason to look closer, never as a finding.

What is the ZeroGPT false positive rate?

Nobody outside the company knows. ZeroGPT advertises "under 1%" on its homepage but names no corpus, no threshold and no independent validation, so the figure cannot be checked. For context, the 14-tool Weber-Wulff study measured a field-wide average of about 2% false positives on human text, rising to about 11% once machine translation was involved, and Stanford measured a 61.3% average on non-native TOEFL essays across seven detectors, ZeroGPT among them. An unsourced sub-1% claim is not a prediction about your writing.

Is ZeroGPT the same as GPTZero?

No. They are separate companies with confusingly similar names. GPTZero (gptzero.me) was built by Edward Tian in 2023, publishes a methodology, names its own weak cases and advertises 99% accuracy with a 10,000-character free box. ZeroGPT (zerogpt.com) is a different product with a 15,000-character free box, a 98.4% claim and no published methodology. If someone quotes "a ZeroGPT result" to you, confirm which URL they used first.

Why did ZeroGPT flag my writing when I wrote it myself?

Most likely because of register rather than anything you did wrong. Classical detectors read perplexity (how predictable your word choices are) and burstiness (how much that varies across the document). Formal academic prose, technical writing, second-language English and heavily edited work all score low on both, which is the same statistical signature AI text produces. The classifier is not wrong about the pattern; it is wrong about what the pattern means.

Does ZeroGPT detect paraphrased or humanised AI text?

Usually not, and this is a limitation of the whole category rather than one product. The Weber-Wulff study measured average accuracy falling to roughly 26% on AI text paraphrased through QuillBot, with about 71% of it going undetected entirely, and found that content obfuscation "significantly worsen[s] the performance of tools." The signal these detectors rely on is exactly what a paraphraser is engineered to disrupt.

Is ZeroGPT accurate on languages other than English?

It claims multi-language support, but there is no published evaluation to check it against, and the independent evidence on translated text is discouraging: the Weber-Wulff set found accuracy of about 75% on machine-translated documents and roughly five times the false positive rate, about 11%, once translation entered the pipeline. Our own detector is English-only and we say so rather than claiming coverage we cannot demonstrate.

What should I use instead of ZeroGPT?

Do not replace one unsourced number with another. Use two detectors built on different signals, prefer tools that publish a methodology and name their own failure cases, and weight the result against version history and drafts. Turnitin scored highest of the 14 tools in the Weber-Wulff study, though individuals cannot buy it. For a second opinion you can run yourself, our free scan takes 1,000 words without an account and we publish our false-positive benchmark with the dataset attached.

Related

More on detector accuracy and false positives.

Further reading

Don't act on a number you can't check. Get a second opinion.

ZeroGPT will not tell you what its 98.4% was measured on. We will: 1,180 academic papers, 5.85% false positives, dataset downloadable. Scan the same passage and compare the two readings yourself, 3 checks a day, no account.

Start free, no card Read our methodology
A benchmark you can audit · Sentence-level highlights · No headline accuracy percentage, on purpose · No signup for the free tier