It can move a score substantially. It cannot make your writing undetectable, and we are not going to print that word about our own product. The one published study we can point to measured GPTZero at 94.1% on unmodified output from DeepSeek's reasoning model and 52% once that same output had been humanised. That is a real, large drop. It is also a coin flip rather than a bypass, and it was measured on one model family at one point in time by researchers who called humanisation "the most effective adversarial attack" they tested.
There is a second thing worth saying immediately, because most pages on this keyword are quiet about it. No third-party tool can show you a GPTZero score. GPTZero does not expose its verdict for other products to display, so any tool implying it is optimising against a live GPTZero reading is showing you something else. We score before and after with our own detector, label it as ours, and tell you what that is worth.
Four things that are supportable from published work, and one claim you will see everywhere that is not.
"Undetectable," "bypass guaranteed," "100% human score." These are claims about a moving target made by the party selling the rewrite, and none of them survive contact with the published record. We will not write them, including about ourselves, and that policy costs us conversions on exactly this keyword. If you want the word "undetectable", plenty of sites will sell it to you. What you would be buying is the word.
There is also a category error worth naming. A humanizer is a rewriting tool. It changes your text. If what you actually need is to show that you wrote something, a rewrite is the wrong instrument entirely, because it replaces the thing you are trying to defend. In that situation read how to prove you did not use AI instead, and keep your version history.
You cannot sensibly rewrite against a detector without knowing what it reads, and GPTZero is unusual in this market for publishing that.
GPTZero's approach centres on two statistical properties. Perplexity is how surprising your next word is to a language model: predictable word choices give low perplexity. Burstiness is how much that predictability varies across the document. Human writing is uneven, with a long qualified clause followed by four blunt words. Model output trends toward a consistent mid-length rhythm at a consistent level of predictability. Low perplexity plus low variance is the signature.
This explains why the tool behaves the way it does on writing nobody generated. Formal academic prose, technical documentation, careful second-language English and heavily copy-edited text all score low on both measures, which is why they get flagged.
To its credit, GPTZero publishes a methodology write-up and names its own weak cases, which most of this market does not. Its stated headline figures are around 99% accuracy, up to 99.9% in some published framings, and 96.5% on mixed human and AI documents. It also states plainly which situations it handles badly: writing by non-native English speakers, passages under about 250 words, and paraphrased text. And it tells institutions in writing that its results "should not be used to punish" students or treated as a final verdict.
The free box accepts up to 10,000 characters, which is roughly 2,300 words, not 10,000 words as it is frequently misreported. Below about 250 words its own documentation describes scoring as noisier, because perplexity and burstiness statistics need length to stabilise. Our full audit of those claims is at is GPTZero accurate.
If you are here because GPTZero flagged writing you produced yourself, note that two of the three failure modes GPTZero itself names may apply to you before any rewriting enters the picture. That is an argument you can make with GPTZero's own documentation, and it is a stronger position than a rewrite.
Sources: gptzero.me published methodology, FAQ and product pages, claims as displayed September 2026. Figures are GPTZero's own, reported here as theirs.
One study gives a clean before-and-after on GPTZero specifically. It is the most useful thing on this page, in both directions.
Alshammari and Rao (University of Missouri, arXiv:2507.17944) tested six detectors over 294 samples, including human question-and-answer text written between 2011 and 2021, before current models existed, alongside model output and deliberately humanised variants. GPTZero's results across those conditions:
| Condition | GPTZero detection |
|---|---|
| DeepSeek-V3 output, unmodified | 100% |
| Human-written text, correctly identified | 98.2% |
| DeepSeek-V3 output, paraphrased | 92.61% |
| DeepThink reasoning output, unmodified | 94.1% |
| DeepThink reasoning output, humanised | 52% |
Read the 52% against the row it actually belongs to. The humanised samples were DeepThink output, so the honest same-generator comparison is 94.1% down to 52%, a fall of roughly 42 points. Setting it against the 100% row instead would mix two different generators and overstate the effect, which is worth avoiding even though it would flatter the case for humanizing. Read the other way, 52% means about half the humanised samples were still caught. The authors described humanisation as the most effective adversarial attack they tested, and that is what "most effective" bought: a coin flip.
Weber-Wulff et al. (2023), 14 tools across 756 test cases, found the same direction at field level. Average accuracy fell from about 74% on raw ChatGPT output to about 42% on manually edited AI text and about 26% on QuillBot-paraphrased text, with 71% of paraphrased content going undetected. The paper concludes that content obfuscation techniques "significantly worsen the performance of tools." It also found the field's dominant failure is false negatives, meaning these tools miss AI more often than they wrongly accuse humans.
These are measurements of particular detector versions against particular generator versions, published in 2023 and 2025. The Missouri study tested DeepSeek, V3 and DeepThink, a generation since superseded by V4 and later releases, and it is the only published DeepSeek-specific detector study. No equivalent measurement exists for current models, and none of these figures were measured on ChatGPT output, so do not read them as GPTZero's behaviour on text from a different lab. Detectors retrain against exactly these attacks, which is why a figure from a paper is not a prediction about your document today. Anyone quoting these numbers as a current success rate, us included, would be overreaching.
Sources: Alshammari and Rao, arXiv:2507.17944; Weber-Wulff et al., International Journal for Educational Integrity 2023, arXiv:2306.15666. TextSight was not evaluated in either study, and we do not construct a comparable figure for our own humanizer by analogy.
This is the part where most tools on this keyword are vague, so here is the mechanism in plain terms.
We score your text before the rewrite, rewrite it, then score the result, and show you both readings with sentence-level highlights. You see which specific lines moved and which did not. In practice the lines that refuse to move are usually the introduction restating the prompt and the conclusion summarising the body, because those are formulaic in genuinely human writing too.
It is our detector's reading, not GPTZero's. GPTZero does not expose its verdict for third parties to display, so nobody can legitimately show you a live GPTZero score inside their own product. A before-and-after against our own model is useful for one specific thing: seeing whether the rewrite changed the statistical properties that this whole family of detectors reads. It is not a prediction about what GPTZero will say, and we would be lying if we framed it as one.
Different detectors disagree with each other constantly. That is the honest reason we show you our own number rather than an implied external one, and it is also the reason a favourable score here should not make you feel finished.
The failure mode nobody advertises is a rewrite that lowers a score by damaging the text. Swap in thesaurus vocabulary, break up every sentence and insert filler hedges, and you will move a detector while producing prose that reads worse and sometimes says something different from what you meant. A rewrite that alters your argument is worse than useless in academic work: it introduces claims you did not make and cannot defend.
So read the output. Check that it still says what you meant, that the citations still attach to the claims they supported, and that no technical term has been paraphrased into something wrong. A humanizer is a drafting aid, not an autopilot.
These are the changes that shift the measurements described above. They are also, not coincidentally, what makes writing better.
This is the single highest-leverage change, because it directly addresses burstiness. Model output settles into a consistent mid-length rhythm. Put a twenty-eight-word sentence next to a five-word one. Let a paragraph end abruptly. The variance is the signal, and it also reads better, which is why this advice is not a trick.
Generated prose generalises because the model is averaging. Named examples, actual figures, a date, a particular objection someone raised, the thing that went wrong in your own attempt: these raise perplexity because they are genuinely unpredictable, and they are also the content that makes an argument worth reading. This is the change that most improves both the score and the work.
"In today's rapidly evolving landscape." "It is important to note that." "In conclusion, this essay has demonstrated." Transitional filler is one of the most predictable things a model produces, and it is dead weight in human writing too. Deleting it raises average perplexity and shortens your draft.
Balanced, hedged, on-the-one-hand prose is low-perplexity almost by construction. A clear claim, with the reasoning for it, is not. If your argument has a view, state it.
Thesaurus substitution produces stilted text and moves scores less than people expect. Inserting deliberate typos degrades your work and is transparent to a reader. And the tricks aimed at watermark-style detection, such as inserting and deleting invisible characters, are aimed at a mechanism that does not apply here, which we explain in our piece on ChatGPT watermark detection.
Stated plainly, in the units they are actually enforced in, because quota copy is where this industry is least honest.
3 rewrites a day, up to 300 words each. Enough to run a paragraph or an abstract and see what the before-and-after looks like on your own writing. No card, no email. Try it at our free humanizer tool.
260 words per rewrite, with 260 words per request. The monthly word budget is deliberate: rewriting generates text through a language model, so the cost profile is different from scanning, and we would rather publish a real limit than advertise "unlimited" and throttle you quietly.
Paid plans start at $9.99/month for 40,000 words a month, or $7.49/month billed annually, and run up through Pro at $19.99/month for 50,000 words. A 10,000-character ceiling applies per request on every tier, including paid ones, so a dissertation gets processed in chunks rather than in one call. Full ladder at pricing.
Chunking is not free of consequences. Rewriting section by section can produce local consistency and global drift, where each part reads well and the whole develops a slightly uneven voice. For anything long, rewrite selectively: the passages that actually read as formulaic, rather than the entire document because the entire document is what you have.
Worth stating directly, because the keyword attracts an assumption about who is searching it.
The largest honest use of a tool like this is a person who wrote something themselves, had it flagged, and wants to understand which sentences are producing the reading. That is a real and common problem, and it falls hardest on the people least equipped to argue about it. Liang et al. (2023) found seven detectors flagging more than 61% of TOEFL essays by non-native English speakers, against near-zero for native-English US eighth-graders. GPTZero was one of the seven, and GPTZero itself names second-language writing as a weak case.
If that is your situation, note the order of operations. Evidence first, rewriting second. Version history, drafts and a conversation about your argument will do more for you than a rewrite, and a rewrite after an allegation can look like exactly the wrong thing. Use the appeal letter template and the process walkthrough first. Use the humanizer on the next draft, while you are still writing it.
The other genuine use is ordinary editing. Seeing which of your paragraphs are formulaic is useful information whether or not anybody is running a detector on you, and the changes that help are the ones listed above: vary the rhythm, add real detail, cut the scaffolding, commit to a claim. That is just writing better, with a measurement attached.
We will not tell you a rewrite makes work undetectable, we will not display a score we did not produce, and we will not describe our own detector's verdict as proof of anything. If your institution prohibits AI assistance, a rewriting tool does not change what you did, and no score from us will resolve that. That is your call to make, with the actual rule in front of you.
Including the ones with answers we would rather not have to give.
No, and we will not use that word. The best-documented figure is from Alshammari and Rao (arXiv:2507.17944), where GPTZero detected 94.1% of unmodified DeepThink reasoning output and 52% of that same output after humanising. That is a large reduction and it is still roughly half the samples caught. Note the comparison has to stay within one generator: the study's 100% figure is for unmodified DeepSeek-V3, a different condition, so quoting 100% against 52% would overstate the drop. Those researchers called humanisation the most effective adversarial attack they tested, so 52% is close to the ceiling that approach reached in published testing, not a floor. Detectors also retrain against exactly these techniques, so any guarantee about a moving target is marketing.
No. GPTZero does not expose its verdict for third-party products to display, so a tool implying it is optimising against a live GPTZero reading is showing you something else, usually its own model relabelled. We score before and after with our own detector and say so on the result. That tells you whether the rewrite changed the statistical properties this whole family of detectors reads, which is useful, but it is not a prediction of what GPTZero will output.
Primarily perplexity, meaning how predictable your word choices are to a language model, and burstiness, meaning how much that predictability varies across a document. Low perplexity with low variance is the AI signature. GPTZero publishes a methodology write-up and names its own weak cases: writing by non-native English speakers, passages under about 250 words, and paraphrased text. Its free input box accepts up to 10,000 characters, roughly 2,300 words, which is often misreported as 10,000 words.
Ours has a free path. 3 rewrites a day at up to 300 words each with no account at all, or 260 words per rewrite on a free account. Paid plans start at $9.99/month for 40,000 words a month, or $7.49/month billed annually, and run up through Pro at $19.99/month for 50,000 words. A 10,000-character per-request ceiling applies on every tier, so long documents get chunked.
It can, and this is the failure mode nobody advertises. A rewrite that lowers a score by swapping in thesaurus vocabulary and inserting hedges can also alter your argument, detach a citation from the claim it supported, or paraphrase a technical term into something wrong. In academic work that is worse than useless, because it introduces claims you did not make and cannot defend. Read the output, check the citations still attach, and treat it as a drafting aid rather than an autopilot.
Probably not first. If the work has already been questioned, a rewrite replaces the thing you are trying to defend and can look like the wrong move. Gather process evidence instead: version history, drafts, notes, and a conversation about your argument. Two of the three weak cases GPTZero itself publishes, second-language English and paraphrased text, may already apply to you, and Liang et al. (2023) measured seven detectors including GPTZero flagging over 61% of TOEFL essays by non-native speakers. Cite that. Use a humanizer on the next draft, while writing it.
Varying sentence length deliberately, which directly addresses burstiness and is the highest-leverage single change. Adding specific verifiable detail, such as names, figures and particular objections, because genuine specificity is unpredictable. Cutting transitional scaffolding like "it is important to note that", which is among the most predictable text a model produces. And committing to a clear position instead of hedging. Thesaurus substitution and deliberate typos move scores less than people expect and damage the writing.
Its published claims audited against the independent record, including the caveats it names itself.
Read the audit →The signals it reads, and how to cross-check a flag rather than argue with the percentage.
Read the guide →The full tool: rewrite, then see the before and after scored at sentence level.
See the tool →What our number actually measures, and why we show sentence highlights instead of one verdict.
Understand the score →If you want a detector with a downloadable false-positive benchmark beside the claim.
See the alternative →The right instrument when the problem is authorship rather than phrasing.
See the evidence list →Rewrite a paragraph and get both readings from our own detector, labelled as ours, with sentence-level highlights showing which lines shifted and which refused to. 3 rewrites a day, 300 words each, no account. We will not print the word undetectable.