Grammarly advertises 99% detection accuracy and a #1 ranking on RAID, an independent academic benchmark of more than 670,000 texts. Both claims are real and checkable, which already puts Grammarly ahead of most of this category. But RAID scores detectors on a specific metric: accuracy at a fixed 5% false positive rate. A top RAID score therefore describes a detector tuned so that one in twenty human-written documents is flagged by design.
There is a second confusion underneath the first. Grammarly ships two different things that people call "the Grammarly AI detector": a probabilistic detection score, and Authorship, which is provenance telemetry recorded while you type. They answer different questions and have completely different accuracy profiles. Below: what the 99% rests on, what Grammarly itself says about its limits (it is unusually candid), where independent testing disagrees, and which of the two tools you are actually being judged by.
Unusually for this category, Grammarly names its benchmark. That makes the claim auditable, which is the whole reason this page can say anything useful.
Grammarly's AI detector page states that it "achieves 99% detection accuracy and ranks #1 on RAID's independent benchmark," and its company blog says the detector "ranked #1 on RAID's leaderboard for AI detection quality." It describes RAID as testing "detection systems under identical conditions using over 670,000 texts across different writing styles, AI models, and even attempts to fool the detectors."
The output is a percentage split between "Resembles AI text" and "No AI patterns found." Grammarly says it identifies output from "Grammarly and other tools like ChatGPT, Gemini, and Claude." Basic scanning is available without signup; the full AI Detector agent comes with Grammarly Pro.
RAID is a genuine academic benchmark, published by Dugan et al. at ACL 2024, and naming it is to Grammarly's credit. But you have to read the metric. RAID tunes each detector's threshold to produce a 5% false positive rate on the human-written portion of the dataset, then reports accuracy on the machine-generated portion.
So "99% accurate, #1 on RAID" unpacks to: when calibrated so that it wrongly flags 5% of human documents, it correctly catches about 99% of machine-written ones on RAID's corpus. That is a strong result. It is not the same sentence as "99% of its verdicts are right," and it is emphatically not "1% of human writing gets flagged."
RAID's designers chose a fixed 5% FPR precisely because false positives are the expensive error, and their own paper concludes that detectors "struggle to operate well at safe false positive rates, regularly struggle to generalize to unseen generators and decoding strategies, and are susceptible to adversarial attacks." The benchmark Grammarly tops is a benchmark whose authors are pessimistic about the field.
Sources: grammarly.com/ai-detector and Grammarly company blog, claims as displayed September 2026. Dugan et al., "RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors," ACL 2024 (arXiv:2405.07940).
If you take one number from this page, take that one. At a 5% false positive rate, a class of 200 students producing one essay each yields roughly ten flagged papers written by humans. That is the cost structure of a top-ranked detector operating at its benchmark threshold, and it is why every serious source, Grammarly included, says the score is a starting point rather than a finding.
Most reviews of "the Grammarly AI detector" silently mix these together. They work on opposite principles and fail in opposite ways.
The AI detector is a classifier. It reads finished text and estimates how much of it resembles AI writing. Like every classifier, it is probabilistic, it can be wrong in both directions, and it has no idea how the document was actually produced.
Authorship is not a detector at all. It records provenance while you write: whether text was typed, pasted, dictated, generated by Grammarly, or copied from somewhere else, and it produces a report of that origin. Grammarly describes it as generating "a comprehensive report detailing the content's origin: human-typed, AI-generated, edited with traditional grammar checking, or sourced externally."
| AI detector | Authorship | |
|---|---|---|
| What it looks at | The finished text | Your editing session as it happens |
| Method | Statistical classifier, probabilistic | Event log, near-deterministic on paste |
| Can it be wrong about clean human writing? | Yes, that is the false positive rate | Rarely on what it can see |
| Blind spot | Register, paraphrasing, short text, ESL prose | Anything written outside Grammarly’s editor |
| Can it be defeated by retyping AI output? | No, it reads the text either way | Yes, typed input looks typed |
| Useful as evidence? | As one signal only | Stronger, but only with prior enrolment |
If you are being assessed on an Authorship report, the argument is about telemetry: what the log recorded, whether you were enrolled, and whether the work happened inside Grammarly at all. Authorship covers only activity in Grammarly's own editor, so a document drafted in Word or Docs without the integration is simply invisible to it. That is a coverage limit, not an accuracy limit.
If you are being assessed on an AI detection percentage, you are arguing about a probability, and every caveat on this page applies. Ask which one it was before anything else. They are not interchangeable and they do not corroborate each other.
Grammarly's published guidance is more careful than most institutions' use of it. These are their words, not ours, and they are worth quoting back.
From Grammarly's own writing on detector accuracy: "No AI detector can conclusively determine whether AI was used to produce text." And: AI detection "is best used as a starting point for review rather than as final proof of authorship," and it "should not be the sole measure of originality." Its detector page repeats that the tool "cannot definitively conclude whether or not AI was used."
Grammarly also names the false positive risk that most of this industry avoids mentioning: "Writers who speak English as an additional language are particularly vulnerable to misclassification because their writing may differ from the datasets used to train AI detectors," and "formal writing that is highly structured or writing from people who speak English as an additional language can sometimes resemble AI-generated patterns."
Grammarly has publicly criticised competitors that "definitively state whether the analyzed content contains AI" as not acting responsibly. On disclosure, they are one of the better actors in this category. The gap is not between Grammarly and honesty; it is between Grammarly's documentation and how institutions consume the number.
Sources: Grammarly, "Are AI Detectors Accurate? What You Should Know" and "From AI Detection to Authorship," grammarly.com, as published September 2026.
Outside RAID, the picture is weaker, and the sources have commercial interests you should weigh. We flag that rather than laundering it.
Several published spot-tests report Grammarly missing a large share of AI text. Pangram Labs, testing 30 detectors on a set of nine AI samples and three human ones, reported Grammarly correctly flagging none of the nine AI samples while correctly clearing all three human ones. Originality.ai's published review of Grammarly reports an F1 of 0.364 and recall of 0.222, which would mean missing roughly 78% of AI text, and titles its review "Easy to Bypass."
Read those with care. Both Pangram and Originality.ai sell competing detectors, and a twelve-sample test is far too small to establish a rate. What is notable is the direction: multiple independent tests find low recall, and low recall is exactly what you would expect from a detector tuned conservatively to protect against false positives. A tool can be simultaneously #1 at a 5% FPR on RAID and permissive in default product use, because those are different operating points.
Reported figures suggest raw ChatGPT output is caught far more often than Claude or Gemini output. That pattern matches the wider literature: RAID's own authors found detectors "regularly struggle to generalize to unseen generators and decoding strategies," and Elkhatat et al. (2023) measured accuracy swings of 15 to 30 percentage points depending on which model produced the text. A detector's headline number is anchored to the generator mix it was calibrated on, and that mix moves every few months.
Sources: Pangram Labs comparative test (2026); Originality.ai, "Grammarly AI Content Detector Review"; Elkhatat et al., "Evaluating the efficacy of AI content detection tools," International Journal for Educational Integrity, 2023. The first two are published by vendors of competing detectors.
Four cases, four different answers. Find yours.
Reasonably reliable, and this is the case every detector handles best. If you pasted 800 words straight out of ChatGPT, expect Grammarly to notice, though independent tests suggest it is more permissive here than its RAID ranking implies.
This is where the 5% bites. Highly structured writing has low perplexity and low burstiness, which is the same statistical signature AI text produces. Grammarly says so itself. A flag on a lab report or a literature review is weak evidence of anything.
Treat any flag as unreliable until corroborated. Grammarly names this group as "particularly vulnerable to misclassification," and the Stanford study behind that concern measured an average false positive rate of 61.3% on non-native TOEFL essays across the seven detectors it tested, against a near-zero rate on essays by native English-speaking US eighth-graders. Grammarly was not one of the seven (they were Originality.AI, Quil.org, Sapling, OpenAI’s classifier, Crossplag, GPTZero and ZeroGPT), so that figure is not a measurement of Grammarly. What carries over is the mechanism, which applies to detectors as a class, ours included.
Expect a miss. The 14-tool Weber-Wulff study measured field-wide accuracy dropping to roughly 26% on QuillBot-paraphrased AI text, with about 71% going undetected, and RAID's authors report detectors are "susceptible to adversarial attacks." A clean score after an editing pass is not a clearance.
Sources: Liang et al., Patterns (Cell Press), 2023; Weber-Wulff et al., International Journal for Educational Integrity, 2023; Dugan et al., ACL 2024.
We sell a detector too, so weigh this accordingly. Here is what we will and will not put a number on.
We do not publish a single headline accuracy percentage. Not 99%, not any figure. One number across every length, genre and register is misleading, and this page has just spent 2,000 words showing why a headline number needs its threshold printed beside it. What we publish instead are bands: on long-form English of roughly 300 words or more, against 15 current commercial model families, our internal benchmark sits at roughly 88 to 92% accuracy; under about 100 words it falls to roughly 70 to 78%. Those are our own internal results, not an independent evaluation, and we label them that way.
On false positives we published the measurement with the data attached: 1,180 academic papers, 5.85% measured false positive rate, downloadable per document. That is in the same range as the 5% threshold RAID uses, and we would rather show you the corpus than advertise a better-sounding number without one.
We have not submitted to RAID. Grammarly has, and that is a real point in their favour that we are not going to talk around. Our detector is English-only, carries the same documented second-language false-positive risk as the rest of the category, and surfaces a low-confidence flag on borderline samples instead of rounding them into a clean verdict. We do not yet have an independent third-party benchmark, we want one, and we will not claim one before it exists.
Sources: our accuracy methodology page and our published false-positive benchmark, both of which link the underlying dataset.
Accurate at a specific operating point, which is the part that gets lost. Grammarly ranks #1 on RAID, a real academic benchmark of over 670,000 texts, with a claimed 99% detection accuracy. But RAID measures accuracy at a fixed 5% false positive rate, meaning the detector is tuned to wrongly flag one in twenty human documents in exchange for catching nearly all machine-written ones. Grammarly itself says no AI detector "can conclusively determine whether AI was used to produce text."
RAID tunes each detector's threshold until it produces a 5% false positive rate on the human-written part of its dataset, then reports how much of the machine-generated part the detector catches at that setting. Ranking first means Grammarly caught the most machine text at that fixed error rate on RAID's corpus. It does not mean 99% of its verdicts are correct, and it does not mean only 1% of human writing gets flagged. RAID's own authors conclude detectors "struggle to operate well at safe false positive rates."
They are different products solving different problems. The AI detector is a statistical classifier that reads finished text and estimates how much resembles AI writing, so it is probabilistic and can be wrong in both directions. Authorship records provenance while you write, logging whether text was typed, pasted, dictated or AI-generated, and produces a report of that origin. Authorship is far more reliable on what it can see, but it only sees activity inside Grammarly's editor, so anything written elsewhere is invisible to it.
Grammarly says it identifies text from "Grammarly and other tools like ChatGPT, Gemini, and Claude." In practice, generalisation across generators is uneven across every detector: published tests report substantially higher catch rates on raw ChatGPT output than on Claude or Gemini, and RAID's authors found detectors "regularly struggle to generalize to unseen generators and decoding strategies." Elkhatat et al. measured accuracy swings of 15 to 30 percentage points depending on the source model.
Using Grammarly's grammar and clarity suggestions is not AI generation, and Authorship reports it as a distinct category: "edited with traditional grammar checking." The AI detector, though, reads the finished text and does not know how it got that way, so heavily revised prose can read as smoother and less bursty than a first draft, which is the same signature it associates with AI. If this matters for an assignment, keep your version history.
It can, and Grammarly says so plainly: writers who "speak English as an additional language are particularly vulnerable to misclassification because their writing may differ from the datasets used to train AI detectors." The underlying research is stark. Stanford measured a 61.3% average false positive rate on non-native TOEFL essays across seven detectors, against a near-zero rate on essays by native English-speaking US eighth-graders. Grammarly was not among those seven, so the number is not a measurement of Grammarly; what transfers is the mechanism, since second-language academic writing has lower perplexity and lower lexical variance, which overlaps the signal every detector in this class reads. This is an industry-wide structural issue, ours included, not a Grammarly-specific defect.
It should not be, and Grammarly agrees. Its published guidance says detection "is best used as a starting point for review rather than as final proof of authorship" and "should not be the sole measure of originality." If a score is being treated as proof, ask for the threshold the institution acts on in writing, re-scan on an independent detector, and produce version history. We keep an appeal-letter template that cites the published literature rather than arguing about your percentage.
What you get if you want AI detection and a rewriter in one place rather than as a Pro add-on.
See the alternative →The same audit applied to the tool most universities actually reference.
Read the audit →A 98.4% claim with no methodology attached, measured against the peer-reviewed record.
Read the audit →What a 5% false positive rate costs a real cohort, and who absorbs it.
Read the guide →Our rate measured at a stated threshold, with the corpus attached so you can check it.
Check our numbers →Quote Grammarly's own caveats back, rather than disputing its percentage.
Use the template →A score tuned to flag one human document in twenty is a starting point, not a finding. Read the same passage against a second detector and keep both results with timestamps. 3 checks a day, no account, no card.