Home · Blog · AI Detection
AI DETECTION

Do AI Humanizers Work? What the Research Actually Shows

Rewriting can change a detector’s verdict — peer-reviewed research says so. But detector accuracy ranges from 33 to 81 percent between products, so no tool can promise a durable result. What humanizers really do, and the cost nobody mentions.

DO

The short answer

It depends entirely on what you mean by "work", and the two meanings pull in opposite directions.

If you mean "will this stop a detector flagging my text", then sometimes, unreliably, and never in a way anyone can promise you in advance. If you mean "will this make flat machine-written prose read like a person wrote it", then yes, and that is a result you can actually check for yourself.

The honest version is that the first question has a well-documented answer in the academic literature, and it is not the answer most humanizer marketing pages imply. Below is what the research says, what it does not say, and what we think the useful test is.

What an AI humanizer actually does

An AI detector does not look up your text in a database. It scores statistical properties of the writing: how predictable each word is given the words before it, how much sentence length varies, how evenly the rhythm is distributed. Raw language-model output tends to sit in a narrow band on all three. It is fluent, evenly paced and slightly too smooth.

A humanizer rewrites text to move it out of that band. It varies sentence length, swaps predictable connectives, breaks up uniform paragraph shapes, and replaces the register that models default to. That is the whole mechanism. There is no cloaking layer and no watermark removal, because for text there is usually no watermark to remove.

We go through the detection side of this in more detail in how AI detectors work.

The evidence that detectors can be defeated

This part is not controversial and it does not come from humanizer vendors. It comes from peer-reviewed work and, in one case, from a learning-platform company that had every commercial reason to conclude the opposite.

  • Paraphrasing alone was enough. Computer scientists at the University of Maryland asked directly whether AI-generated text can be reliably detected, and concluded it cannot, with simple paraphrasing sufficient to evade the detectors they tested (Sadasivan et al., 2023).
  • Detector accuracy varies enormously. A study of 14 AI detectors run by academic researchers across six countries found accuracy ranging from just 33 to 81 percent depending on the provider and the method (Weber-Wulff et al., 2023).
  • A major LMS vendor tested detection and walked away. Anthology, which owns Blackboard, ran a beta test across May and June 2023 with 65 client institutions submitting more than 1,000 texts, some authentic and some AI-generated, specifically to evaluate putting AI detection inside SafeAssign. 80 percent of respondents felt the detectors were, at best, only able to "sometimes" identify texts correctly, and Anthology and its participating clients concluded that AI detection "is not currently fit for purpose in education" (Anthology white paper). We cover what that means for Blackboard users in does Blackboard detect AI and the SafeAssign AI checker explained.

So the narrow claim "rewriting can change a detector's verdict" is true, and we are not going to pretend otherwise just because we also sell a detector.

Why "it beat the detector" is a much weaker result than it sounds

Here is the part the marketing pages leave out. A passing score is a measurement against one model, at one moment, on one piece of text. It generalises badly, for four reasons.

Detectors disagree with each other. That 33 to 81 percent range is not noise, it is the spread between products. Text that clears one detector can be flagged by the next, and you usually do not get to choose which one your reader runs.

The models move. Detector vendors retrain. A result from last month is not a guarantee about this month, and nothing about a past score carries forward.

A score was never a verdict. Every serious detector, ours included, outputs a probability that text resembles machine-generated writing. It is not proof of authorship in either direction. A low score does not certify that a human wrote something, which is exactly why a "passing" result proves less than it appears to. See can AI detectors be wrong.

Detectors are sensitive to surface features that have nothing to do with who wrote the text. This cuts both ways and it is the strongest argument against treating any single score as meaningful. It is also why false positives land hardest on people who did nothing wrong: Stanford researchers found that detectors misflagged 61.3 percent of TOEFL essays written by non-native English speakers (Liang et al., 2023). More on that in AI detector false positives.

The cost nobody puts on the sales page: meaning drift

A humanizer rewrites your sentences. Rewriting changes meaning, and the more aggressively a tool rewrites to move a score, the more it changes.

The things that break first are the things that matter most in serious writing. Numbers get rounded or transposed. Hedges disappear, so "the data suggests" becomes "the data proves", which is a different and often false claim. Technical terms get swapped for near-synonyms that mean something else in the field. Citations drift away from what the cited source actually said.

For a blog intro this is survivable. For a literature review, a clinical summary, a legal note or a financial explainer it is the actual risk, and it is far more likely to cause you a problem than a detector score ever was. If you use any rewriting tool on work that has to be correct, diff the output against your original and read the changes. Do not skim them.

Where humanizers genuinely help

Strip away the detector framing and there is a real editing problem underneath.

Raw model output has recognisable habits: sentences that are all roughly the same length, a fondness for three-item lists, connectives that arrive exactly where you expect them, and a hedged, agreeable register that never commits to anything. That prose is not flagged because a machine wrote it. It is flagged because it is flat, and flatness is measurable.

Fixing that is ordinary editing, and a good rewrite does it faster than you would by hand. The output is better to read whether or not anyone ever runs a detector over it. That is the case for these tools that survives contact with the evidence, and it is the one we are comfortable making for our own humanizer.

Where they do not help, and where we will not help

If the goal is to submit machine-written work as your own in a course that forbids it, a humanizer is not a solution to that problem. It is a way of making the problem harder to see, the underlying integrity issue is unchanged, and the tool cannot promise you the outcome you are buying it for.

We build detection and writing-trust tools. We do not publish evasion guides, we do not tell anyone how to tune text against a specific detector, and we are not going to start. If you have been flagged for work you actually wrote, the useful path is evidence rather than rewriting: draft history, version control in Google Docs or Word, notes and outlines, and a conversation about the work itself. We set that out in how to prove you did not use AI and how to write an AI detection appeal letter.

How to judge a humanizer honestly

Four tests, in the order they matter.

  1. Does it preserve meaning? Run a paragraph you know cold through it and read the output against the original line by line. Check every number, every citation and every hedge. This is the test most tools fail and almost nobody runs.
  2. Does it still sound like you? A rewrite that reads as generically human has swapped one borrowed voice for another. If the point was for the writing to be yours, check that it still is.
  3. Is it readable prose, not scrambled prose? Some tools move a score by introducing odd word choices and broken rhythm. That lowers a detector number and makes the writing worse, which is the opposite of the thing you wanted.
  4. Does the vendor promise "undetectable"? Treat that as a red flag rather than a feature. Given how much detector accuracy varies between products and how often they retrain, it is not a promise anyone is in a position to keep.

What we claim, and what we do not

TextSight ships a humanizer and an AI detector, so we have an obvious interest here and you should read this section with that in mind.

We do not claim our humanizer makes text undetectable. We would have to be able to speak for every detector on the market, including ones that do not exist yet, and we cannot. What we do claim is narrower: it rewrites flat machine prose into something that reads better, it shows you a humanization score so the change is visible rather than asserted, and it is built to preserve your meaning while doing it.

On the detection side we publish our limits rather than a headline accuracy number, including that the detector is English-only, in AI detection limitations and our accuracy methodology. A score from us is a signal worth looking at. It is not a verdict, and we say so on the page where you get it.

Try it on your own writing

DB

Founder & CEO · TextSight

Writing about AI detection, humanization, and the strange new craft of writing in 2026. Operates Lacewing Technologies from Maharashtra, India.

Try the detector free.

Paste any text. See where AI signals show up. Fix what's flagged in minutes.

Start free — no card More from the blog