HomeVoice Detection › Accuracy

How accurate is AI voice detection, really?

Short answer: it depends so heavily on the audio that a single percentage would tell you almost nothing, which is why we do not publish one. Long answer below — what actually moves reliability up and down, what a false positive looks like and who it lands on, and how to weigh a result before you act on it. This is the page we would want to read before trusting anyone else's detector.

Try the voice detector The technical version
Our position

Why we do not publish an accuracy percentage.

This is a deliberate choice, and it costs us conversions against competitors who do. Here is the reasoning.

An accuracy figure is a measurement taken under conditions. Change the conditions and the figure changes — not slightly, but by margins large enough to reverse the conclusion you would draw from it. In voice detection the conditions that matter are which generator produced the audio, how the audio reached you, how long it is, and how much processing sits between the microphone and your file. None of those are constant, and none of them are things we know about your clip.

So "99% accurate" answers a question nobody is asking. The question you actually have is "how much should I trust this result, on this file" — and a headline number cannot answer it. Worse, it actively interferes, because it invites you to apply a laboratory figure to a forwarded voice note where it does not remotely hold.

What to ask any vendor quoting a number

  • Measured on which dataset? In-house test sets tend to flatter. Public benchmarks are comparable.
  • Against which generators, and were any unseen during training? Performance on known generators says little about the next model release.
  • Under what audio conditions? Clean studio audio, or phone-codec audio and re-encoded voice notes?
  • At what false-positive rate? A detector that calls everything synthetic catches 100% of deepfakes. The error rate on genuine recordings is the half that gets left out.
  • Measured when? A figure from a generation of models ago describes a world that no longer exists.
What we will commit to. Everything below about what raises and lowers reliability, stated plainly. A confidence value shown with every result rather than a bare verdict. No claim that any result is proof. And no accuracy percentage anywhere on this site that is not tied to a named benchmark and stated conditions — if we publish one, it will come with all five answers above.
In practice

What actually moves the reliability of a result.

These apply to every detector in this category, ours included. Use them to judge your own clip before you judge the answer.

↑ Makes the result more reliable

  • Thirty seconds or more of continuous speech
  • A single speaker, no crosstalk
  • Lossless or high-bitrate audio — WAV, FLAC, or a good MP3
  • The original file, not a forwarded or re-uploaded copy
  • Quiet background, no music underneath
  • A known-genuine control recording to compare against

↓ Makes the result less reliable

  • Clips of a few seconds
  • Phone-call audio and messaging-app voice notes
  • Anything re-recorded through a speaker
  • Traffic, crowd or music in the background
  • Heavy studio processing — noise reduction, de-essing, pitch correction
  • Clips assembled from more than one source
  • Audio from a generator newer than the detector

The last item on the right deserves emphasis, because it is structural rather than incidental. Detection models learn the traces left by the synthesis systems they were trained against. Generative speech models are released faster than detection models are retrained. Every deployed detector is therefore permanently somewhat behind the newest generators, and no amount of care with your audio fixes that. It is the honest ceiling on this whole category.

The error that matters

False positives, and who pays for them.

A missed deepfake is a failure. A wrongly flagged real person is a different kind of failure, and it is the one worth designing against.

A false positive is a genuine human recording scored as synthetic. The asymmetry is the point: when a deepfake slips through, you are back where you started. When a real recording is flagged, someone now has to prove a negative about their own voice — to an employer, a platform, a client, or a family member who has already decided.

What tends to get wrongly flagged

SituationWhy it reads as synthetic
Professionally produced audioNoise reduction, compression, de-essing and EQ remove exactly the untidiness that marks a recording as physical. A polished podcast can look cleaner than reality allows.
Treated rooms and close micingA voice booth has almost no room tone. The absent noise floor is a genuine property of a good recording space and also a synthesis tell.
Very short clipsNot enough speech to establish a pattern, so the score drifts toward whichever side the audio conditions happen to favour.
Even, measured deliverySomeone reading prepared remarks calmly and clearly produces regular prosody. Regularity is one of the things the detector is looking for.
Heavily transcoded uploadsAudio that has been through several platforms carries stacked codec artefacts that a model has no principled way to distinguish from generation artefacts.

Notice the pattern: several of these describe people who record well. Voice actors, podcasters, broadcasters, anyone using a decent microphone in a treated space — professionals are more exposed to false positives than someone talking into a phone in a kitchen. That is a real fairness problem in this category and it deserves saying out loud rather than being buried.

If a genuine recording of yours was flagged: submit a longer, less-processed take if one exists; submit the rawest file you have rather than the published master; and ask whoever is relying on the result what their false-positive rate is and how a person is supposed to contest it. A process with no route to challenge a flag is not a process.
Decision guide

How to read your result.

ResultReasonable readingNot a reasonable reading
Confidently synthetic, clean audio Strong reason to treat the recording as untrusted and to investigate its provenance seriously. Proof that a named person faked it, or grounds on its own for a disciplinary or legal action.
Confidently synthetic, compressed audio Worth acting on cautiously — compression usually pushes scores toward uncertainty, so a confident result despite it is meaningful. Equivalent to the clean-audio case. The evidence base is thinner.
Confidently human The audio carries the physical signature of a real capture. That the recording is honest, unedited, in context, or of the person you think it is. Detection answers none of those.
Uncertain The file did not carry enough usable signal. Almost always a statement about the audio, not about the voice. Weak support for whatever you already believed. This is the most common misreading and the most costly.

What to do when it is uncertain

  1. Find a better copy. Go back toward the original — fewer forwards, no re-encoding, longer duration.
  2. Run a control. A recording of the same speaker you are confident is genuine, captured similarly. The comparison method partly cancels out channel effects and is the highest-value move available to you.
  3. Stop relying on the audio. Contact the person through a channel you trust. Check whether anything else from the same occasion exists. Ask who had the file before you.
  4. Accept an open question. Sometimes the honest answer is that the recording cannot be resolved from the audio alone, and deciding anyway is worse than not deciding.
Scope

What this is not for.

Not forensic-grade. A TextSight voice result must never be the sole basis for a disciplinary, employment, financial or legal decision. It is a triage signal that tells you a recording deserves attention.

Concretely, the things it should not decide by itself:

  • Firing or disciplining someone over a recording they say is not theirs.
  • Rejecting a candidate because a screening call scored oddly.
  • Publishing an accusation that a public figure's recording is fabricated.
  • Denying a claim or a transaction on a score alone.
  • Concluding a family emergency is fake — verify by calling back on a number you look up yourself, which works regardless of what any detector said.

In each case the detector's proper role is to tell you where to spend your attention. Where the stakes genuinely justify it, a qualified forensic audio examiner can do what an automated score cannot: examine chain of custody, container structure and edit points, and defend a stated methodology with stated uncertainty.

Two more things we do not claim

We do not name the generator. Attributing audio to ElevenLabs, OpenAI or any specific vendor is a much harder problem than telling human from synthetic, and our detector does not return it. We would rather show nothing than a plausible guess.

We do not claim per-language parity. Detection works on acoustics rather than language, so it is not English-only, but we have not published measured figures per language and will not imply an equivalence we have not tested. The Hindi guide says the same thing in Hindi.

FAQ

Accuracy questions.

How accurate is the TextSight AI voice detector?
We do not publish a single accuracy percentage, because one number cannot describe performance that varies enormously with the generator, the audio conditions and the length of the clip. What we will say is what changes reliability in each direction, and where the result should not be relied on at all. A figure without those conditions attached would be marketing rather than information.
Why do competitors advertise 99% accuracy?
Usually because a number measured under favourable, in-house conditions is technically defensible and commercially useful. It is normally measured on clean audio from generators the detector already knows, and it rarely comes with a false-positive rate. Ask which dataset, which generators, what audio conditions, and what proportion of genuine recordings were wrongly flagged. Numbers that survive those questions are rare.
What is a false positive in voice detection?
A false positive is a genuine human recording that the detector scores as synthetic. It happens most often with heavy studio processing, aggressive noise reduction, very clean booth recordings with an unusually low noise floor, and short clips. It is the error type that does real harm, because it puts a person in the position of disproving something.
What makes a voice detection result more reliable?
Longer continuous speech from a single speaker, minimal compression, no re-recording, low background noise, and a copy as close to the original as possible. A thirty-second uncompressed WAV of one person talking is a far better input than five minutes of a forwarded, re-encoded phone recording.
Does compression really change the answer?
Substantially. Phone calls and messaging-app voice notes use codecs designed to discard the parts of the signal a listener will not consciously miss, which overlaps heavily with the fine detail detection depends on. Expect lower confidence on compressed audio, and read an ambiguous result as genuinely ambiguous rather than as a soft yes.
Can I use a voice detection result to accuse someone?
No. It is not forensic-grade analysis and should never be the sole basis for a disciplinary, employment, financial or legal decision. Use it to decide what deserves investigation, then investigate — provenance, corroboration, and asking the person directly. If the stakes justify it, engage a qualified forensic audio examiner.
Is it less accurate for languages other than English?
The analysis works on acoustic characteristics rather than on language, so it is not restricted to English. We have not published separate measured figures per language and will not imply parity we have not measured. Treat non-English results with the same caution as English ones, and lean on the comparison method where a language-specific reference recording is available.
What should I do with an uncertain result?
Treat it as no information rather than as weak evidence for whatever you already suspected. Try a longer or less-processed copy of the same audio, run a known-genuine recording of the same speaker as a control, and fall back on verification that does not depend on the audio at all — contacting the person through a channel you trust.
Related

More voice detection guides.

A signal worth having. Never a verdict.

Three checks a day, free, no signup. Your audio is never stored.

Try the voice detector All voice guides
No accuracy percentage without a named benchmark and stated conditions