Home · Blog · Academic Integrity
ACADEMIC INTEGRITY

What Percentage of AI Is Acceptable? The Honest Answer

There is no universal acceptable AI percentage. Only one vendor publishes a threshold, it is 20%, and it is advice about when not to act rather than a pass mark.

WH

The question almost always arrives in the same shape: someone has a number, and they want to know whether it is safe. Is 15% fine? Is 30% a problem? Is 8% clean?

Here is the honest answer, and it is worth more than a made-up figure: there is no universal acceptable percentage, because no authority publishes one. No regulator requires AI detection. No accreditation body sets a threshold. No sector-wide standard exists at which a score becomes an allegation. Any article confidently telling you "stay under 20%" is describing one vendor's internal guidance to staff and quietly presenting it as a rule that applies to you.

What does exist is more useful than a pass mark, once you understand it. Let us go through what the number actually is, the one published threshold anybody has, and why aiming at a target is the wrong strategy even when you know it.

The only published threshold, and what it actually says

Turnitin, which is the tool most institutions use because it was already installed for similarity checking, publishes a figure: it recommends that scores below 20% should not be acted upon.

Read that carefully, because it is routinely inverted. It is not a statement that 19% is acceptable conduct. It is Turnitin telling the people who read its reports that below that line its own confidence is too low to support a decision. It is a floor on the tool's reliability, not a ceiling on your AI use.

The reason that floor exists is the false positive rate. Turnitin's own documentation puts it at under 1% at document level above the 20% threshold, and about 4% across the full distribution, measured on an internal evaluation set whose size is not disclosed. Below 20% the proportion of flags that are simply wrong climbs to the point where acting on them would produce more injustice than it prevents.

So the only published number in this entire subject is a caution aimed at your marker, not a budget aimed at you.

What the percentage is actually measuring

This is the part that reframes the question. A detector's percentage is not a measurement of how much AI you used. It cannot be, because the detector was not present while you wrote and has no record of your process. It is an estimate of how much of your text carries the statistical fingerprint the model associates with generated writing.

Those are very different things, and they come apart in both directions:

  • Text you wrote yourself can score high. Formal academic register is careful, conventional, hedged and structurally predictable, which is the same low-perplexity profile generated text has. Heavily edited prose scores higher than a rough draft. So does careful second-language English.
  • Text a model wrote can score low. Paraphrasing collapses detection. The largest academic comparison of detectors, Weber-Wulff et al. (2023), measured average accuracy falling to roughly 26% on AI text paraphrased through QuillBot, with about 71% of it going entirely undetected.

A number that can be high for honest work and low for generated work is not a dial measuring your conduct. Treating it as one is the central mistake behind most of the distress on this topic.

The same study concluded that the tools it tested "are neither accurate nor reliable" and that the field's dominant bias is toward classifying output as human-written. The aggregate failure of AI detection is missing AI, not inventing it, which is worth knowing if you have been told the tools are ruthless.

The base rates, from the vendor's own count

Turnitin has published figures that give a sense of what normal looks like at scale. In a release dated 9 April 2024, covering data to 21 March 2024, it reported that since launching AI writing detection in April 2023:

  • over 200 million papers had been reviewed
  • over 22 million, about 11%, showed at least 20% AI writing present
  • over 6 million, about 3%, showed at least 80%

Those are the vendor's own unaudited numbers, and they describe how the tool scored documents rather than how many students used AI, since they include whatever false positives the tool produces. But they tell you something practical: a high score is uncommon without being rare. Roughly one paper in nine crosses the threshold at which Turnitin says a score is worth considering at all.

Why chasing a target number is the wrong strategy

Suppose you knew your institution's threshold exactly. Optimising against it still fails, for three reasons.

Different detectors disagree constantly. A score is a property of one model's opinion, not of your document. Getting a favourable reading from a free checker tells you very little about what your institution's tool will output, because they were trained differently on different data. Several universities have looked at this and stepped back entirely: Vanderbilt disabled Turnitin's AI detection in August 2023 and published its reasoning, which centred on not being able to see how the determination was made. Its arithmetic is worth keeping in mind, roughly 75,000 papers a year, so even a 1% false positive rate implies about 750 wrongly flagged papers.

The risk is not evenly distributed. If English is not your first language, you carry more of it, through no fault of your own. Liang et al. (2023), published in Patterns, ran seven detectors over TOEFL essays by non-native English speakers and found they flagged more than 61% on average, against near-zero false positives on essays by native-English US eighth-graders. No amount of threshold knowledge protects you from a systematic bias in what the tool measures.

Editing to lower a score often makes the writing worse. The changes that move a detector most are the ones that increase variation and specificity. Done well, that genuinely improves prose. Done as score management, it produces thesaurus vocabulary, broken-up sentences and inserted hedges, and in academic work it can detach a citation from the claim it supported or paraphrase a technical term into something wrong. A rewrite that alters your argument introduces claims you did not make and cannot defend.

What to do instead

Read your actual policy. Your institution's academic integrity policy and your module handbook are the only authoritative answer for your situation, and they vary by department. Many now state plainly what AI assistance is and is not permitted, and that stated rule governs rather than any percentage. If a declaration is requested, answer it accurately. An accurate declaration of light assistance is a much stronger position than an inaccurate denial, and it removes the detector from the conversation entirely.

Keep your drafting history. This is the single most valuable thing you can do, it costs nothing, and it is better evidence than any score. A Google Docs or Word version history showing work composed over days, with false starts and reordered paragraphs, demonstrates authorship in a way no percentage can. Some tools now read this directly instead of scoring text at all.

Use a detector for editing, not for certification. Running your own draft is genuinely useful for one thing: seeing which passages read as formulaic. In practice the flags on honest work land on the introduction restating the question and the conclusion summarising the body, which is useful information about your writing. It is not a clearance certificate, and no honest vendor should sell it to you as one.

If you have been flagged, argue the method rather than the verdict. Disputing whether you personally are the 4% the tool got wrong is an argument you cannot win by assertion. The published false positive rates, the non-native English findings, and the vendor's own guidance about thresholds are the lever.

Frequently asked questions

Is 20% AI acceptable?

Not as a rule, because 20% is not a pass mark. It is the level below which Turnitin recommends its own scores not be acted on, because its false positive rate climbs below that line. Whether any AI assistance is acceptable in your work is set by your institution's policy, not by a detector's threshold, and some policies permit none while others permit a great deal with declaration.

Is there an acceptable AI percentage for university work?

No universal one exists. No regulator requires AI detection, no accreditation body sets a threshold, and policy varies between institutions and often between departments. The only authoritative source for your situation is your own academic integrity policy and module handbook.

Does a 0% AI score mean my work is safe?

It means one detector, on one reading, did not find the pattern it looks for. It is not a guarantee, because different detectors disagree constantly and a favourable result from one tool does not predict another's output. It also says nothing about plagiarism, which is a separate check answering a different question.

Why does my own writing get a high AI percentage?

Usually because of register rather than anything you did wrong. Detectors read perplexity, meaning how predictable your word choices are, and burstiness, meaning how much that varies across a document. Formal academic prose, technical writing, careful second-language English and heavily edited work all score low on both, which is the same statistical signature generated text produces.

How many papers actually get flagged?

Turnitin reported that of over 200 million papers reviewed between April 2023 and March 2024, about 11% showed at least 20% AI writing present and about 3% showed at least 80%. Those are the vendor's own unaudited figures, they describe how the tool scored documents rather than student behaviour, and they include false positives.

Can a percentage prove I used AI?

No, and better-run institutions say so in their own guidance. A score is what prompts someone to look. What decides real cases is version history, drafts, consistency with your other work, and a conversation about your argument. Someone who wrote a piece can almost always discuss why they structured it as they did.

Try it on your own writing

DB

Founder & CEO · TextSight

Writing about AI detection, humanization, and the strange new craft of writing in 2026. Operates Lacewing Technologies from Maharashtra, India.

Try the detector free.

Paste any text. See where AI signals show up. Fix what's flagged in minutes.

Start free — no card More from the blog