HomeVoice Detection › Choosing a Detector

How to choose an AI voice detector.

Every vendor in this category advertises a number in the high nineties, and almost none of them publish the conditions that number was measured under. This is a buyer's guide rather than a ranking: the questions that separate a real capability from a marketing claim, the answers that should worry you, and a test you can run yourself in ten minutes that will tell you more than any comparison table.

Check an audio clip free The ten-minute test ↓
No ranked list Run the test yourself
Upfront

Why there is no ranked list on this page.

It would be the easiest page on this site to write, and we would not stand behind a word of it.

A “best AI voice detectors” article written without testing the tools is an exercise in restating vendor marketing in a new order. The numbers are unverifiable, the feature tables come from pricing pages, and the ranking usually reflects who has an affiliate programme. We sell a voice detector, which makes us the last people whose unverified ranking you should trust.

A comparison worth publishing would require building a test corpus of genuine and synthetic audio across multiple generators, degrading it through realistic channels, running every tool, and reporting false-positive rates alongside detection rates — including our own results, unflattering ones included. That is real work, it is on our list, and until it is done this page gives you something more useful: the method to evaluate any of them yourself, us included.

The short version: ignore headline accuracy numbers entirely, ask the five questions in the next section, and run the ten-minute test below on your own audio. A vendor's behaviour when asked precise questions tells you more than any figure they publish.
Questions

Five questions that separate capability from marketing.

Ask them of every vendor, including us. The quality of the answer is the signal.

  1. “What is your false-positive rate, and on what?” A detector that flags everything catches every deepfake. The rate at which genuine recordings are wrongly flagged is the half that gets omitted, and it is the error that harms real people. If they cannot answer, the headline number is meaningless.
  2. “Which generators was it evaluated against, and were any unseen in training?” Performance on known generators predicts very little about the next model release. Unseen-attack performance is what predicts real-world behaviour. A vendor who understands the question is already ahead of most.
  3. “Under what audio conditions was that measured?” Clean studio audio, or telephone codecs and re-encoded messaging-app voice notes? If your audio is compressed phone recordings, laboratory figures on pristine files tell you nothing about your case.
  4. “Does it name the generating tool, and how?” Source attribution is much harder than binary detection. A tool that confidently names a vendor should be able to explain what that rests on and how it handles a model released last month. Confident attribution with no explanation is a red flag, not a feature.
  5. “What do your terms say about relying on the result?” Read the actual terms. Many vendors advertise near-certainty in marketing and disclaim all reliance in their contract. That gap tells you what they really believe.

Answers that should worry you

ClaimWhy it should give you pause
“99.9% accurate” with no conditionsNot a claim about your audio. Ask what corpus, which generators, what false-positive rate, measured when.
“Detects all AI voices”Not achievable. Detection degrades against generators it has not seen, and generation outpaces detection by construction.
“Court-admissible” / “forensic-grade”Forensic audio authentication involves chain of custody, provenance and an expert who can be cross-examined. An automated score is not that.
“Real-time protection on your calls”Detection runs on recorded files. Ask precisely what runs live, where, and on what audio.
“Identifies the exact AI model used”Attribution is a research problem. Ask how it performs on a generator released after their training set closed.
A binary verdict with no confidence valueThe underlying quantity is continuous and the threshold is a policy choice. Hiding it hides how much to trust the answer.
Test

The ten-minute test you can run yourself.

This will tell you more about a detector than any comparison table, including ours.

  1. Record thirty seconds of yourself on your phone. Ordinary conditions, no processing. This is your genuine control.
  2. Generate thirty seconds with any TTS tool, including a free tier. This is your synthetic control.
  3. Degrade both identically. Send each through a messaging app, or export both to a low bitrate. Now they resemble the audio you will actually deal with rather than laboratory files.
  4. Run all four files — both originals and both degraded copies — through each detector you are evaluating.
  5. Read the pattern, not the scores. The questions that matter: does it flag your genuine recording? Does it still catch the synthetic one after degradation? Does confidence drop honestly on the degraded pair, or does it stay implausibly high?
What a good result looks like. Confident and correct on the clean pair. Still correct but less confident on the degraded pair. A tool that reports the same high confidence on a mangled file as on a pristine one is not measuring the audio — it is performing certainty, and that is worse than being wrong occasionally.

Beyond accuracy

  • What happens to your audio. Stored, or processed and discarded? Used for training? Ours is processed and discarded, never used for training, never shared — check what any other vendor says, in the privacy policy rather than the marketing.
  • Is the free tier enough to evaluate properly? You need several checks to run the test above. Ours gives three a day with no signup, which is enough.
  • Formats and limits. Ours takes MP3, WAV, M4A, OGG and FLAC up to 10 MB, and does not accept Opus directly — which matters if your audio is WhatsApp voice notes.
  • Does it explain uncertainty? A tool that tells you when a file was too short or too degraded to judge is being more useful than one that always produces a confident answer.
Not forensic-grade. A TextSight voice result is a triage signal, not proof, and it must never be the sole basis for a disciplinary, employment, financial or legal decision. Our full position on what a result is worth is here.
FAQ

Buyer's guide questions.

Which AI voice detector is the most accurate?
We do not publish a ranking, because we have not run a controlled comparison and restating vendor marketing in a new order would not be information. Run the ten-minute test on this page with your own audio instead — it answers the question for your specific case, which a general ranking cannot.
Why does TextSight not publish an accuracy percentage?
Because one number cannot describe performance that varies enormously with the generator, the audio conditions and the clip length. A figure without those conditions attached invites you to apply a laboratory result to a forwarded voice note where it does not hold. If we publish one, it will name the benchmark, the generators, the conditions and the false-positive rate.
What is the single most important question to ask a vendor?
What is your false-positive rate, and measured on what. A detector that flags everything catches every deepfake, so the rate at which genuine recordings are wrongly flagged is the half of the picture that determines whether the tool is safe to act on.
Is a higher advertised accuracy always better?
No, and it is often a sign of a less careful vendor. Accuracy figures are trivially inflated by testing on clean audio from generators already seen in training and omitting the false-positive rate. The vendors who answer precise questions precisely tend to be the more capable ones.
Should I use more than one detector?
For anything consequential, yes. Independent signals that agree are stronger than one signal, and disagreement is itself informative — it usually means the audio is genuinely ambiguous. Just do not treat two agreeing tools as proof; they may share the same blind spots.
How do I test a detector without a real deepfake?
Generate one. Any text-to-speech free tier gives you a synthetic sample in a minute, and a phone recording of yourself gives you the genuine control. Degrading both the same way is the step that makes the test resemble reality.
Related

More voice detection guides.

Test us the same way. That's the point.

Three checks a day, free, no signup. Your audio is never stored.

Check an audio clip All voice guides
Three checks a day, free, no signup — enough to run the test on this page