Home › Voice Detection

AI voice detection, and what to do when a recording sounds wrong.

Cloned voices are now good enough to fool the ear, and a few seconds of someone speaking is enough raw material to build one. This is the practical guide: how detection actually works, what a score is worth, what it is not worth, and how to handle the specific situations people land here with — a WhatsApp voice note, a scam call, a Hindi recording, a clip you need to verify before you publish it.

Check an audio clip free How accurate is it?
3 checks a day, no signup Audio never stored
The basics

What AI voice detection actually does.

It reads the recording, not the words. A detector never considers what was said or whether the claim is plausible — it looks at how the sound was produced.

A human voice reaches a file through a long, messy physical chain. Air moves through a throat and mouth, bounces around a room, hits a microphone diaphragm, and passes through a preamp and a codec. Every stage adds irregularity: room reflections, breath, lip and tongue noise, tiny timing drift, a noise floor that never quite goes silent.

Synthetic speech skips almost all of that. A text-to-speech or voice-cloning model reconstructs a waveform directly from a learned statistical model of what speech looks like. It is very good at reproducing the parts of a voice we consciously notice — timbre, accent, intonation — and much weaker at reproducing the parts we do not, because those parts were never the point of the training objective.

Detection lives in that gap. The result is a probability, not a fact, and it is a probability about how the audio was produced — never about whether the person in the recording is trustworthy or whether what they said is true.

The signals a detector weighs

SignalWhat a real recording looks likeWhat often gives synthesis away
Spectral detailEnergy scattered untidily across the frequency range, including high frequencies the codec barely preservesSmoother, more regular spectral shape; sometimes a hard ceiling where the vocoder stops generating detail
Micro-timingPauses between words vary constantly, even mid-sentenceGaps that land in a narrower, more uniform range than a person can produce
Breath and mouth noiseAudible inhales, lip smacks, swallows, the odd stumbleBreath either absent or placed identically each time it appears
Noise floorContinuous room tone that persists through silencesSilences that are cleaner than the room ever was, or noise that starts and stops with speech
Energy contourVolume rises and falls unevenly inside each syllableEnvelopes that repeat a learned shape across different words

No single one of these is decisive. A close-miced studio recording in a treated room can look suspiciously clean; a synthetic clip played through a phone speaker and re-recorded picks up genuine room noise that masks the tells. The verdict comes from weighing them together, which is exactly why it arrives as a confidence score rather than a yes or a no.

Read this before you act

What a score is worth — and what it is not.

Most of the harm in this category comes from people treating a detector output as proof. It is not proof, and every honest vendor in the space will tell you the same thing.

Not forensic-grade. A TextSight voice result is not admissible forensic analysis and must never be the sole basis for a disciplinary, employment, financial or legal decision. It tells you a recording is worth investigating. The investigation is still yours to do.

Conditions that lower the value of any result

These apply to every detector on the market, ours included. If two or more are true of your clip, weight the result much less heavily:

  • Very short audio. Under a few seconds there is not enough speech to establish a pattern in either direction.
  • Heavy compression. Phone calls, voice notes and social-media re-uploads strip out exactly the high-frequency detail detectors lean on.
  • Re-recording. Audio played through a speaker and captured again picks up genuine room acoustics, which can make synthetic speech look human.
  • Background noise. Traffic, music or crowd noise buries the micro-signals.
  • Editing. Clips cut together from multiple sources produce a mixed result that reflects neither source cleanly.

We do not publish a single accuracy percentage

You will see competitors advertise numbers like "99% accurate". Treat those with suspicion unless they name the benchmark, the generators tested, the audio conditions and the false-positive rate. Detection accuracy is not one number — it varies enormously with how the audio reached you and whether the generating model resembles anything the detector was trained against. We explain our position on this in full, including why a number without conditions attached is worse than no number.

We do not guess which tool made the voice

Naming a specific generator — ElevenLabs, OpenAI, Murf — from audio alone is a much harder problem than telling human from synthetic, and our detector does not return it. Rather than display a plausible-sounding guess, we leave it out. If you see a tool confidently attributing a clip to a named vendor, ask what it is basing that on.

Practical

How to check a clip properly.

  1. Get the most original copy you can find. Every forward, download and re-upload degrades the audio and the result. If the clip reached you through three people, ask the first person for their copy.
  2. Check the format and size. The detector takes MP3, WAV, M4A, OGG and FLAC up to 10 MB. If you have a WhatsApp voice note, it is an Opus file and needs converting first — the WhatsApp guide covers exactly how.
  3. Trim to the clearest continuous speech. Thirty seconds of one person talking cleanly is worth more than five minutes of a noisy multi-speaker call.
  4. Run the check and read the confidence, not just the verdict. A low-confidence "likely AI" and a high-confidence "likely AI" are completely different findings.
  5. Corroborate before you act. Where did the file come from? Does a known-real recording of the same person score differently? Comparing against a reference clip is the single most useful thing you can do.
Free, no signup: 3 checks a day. A free account raises it to 5 a day, Starter 25, Pro 100, and Business and Enterprise are unlimited. Audio is processed and discarded — never stored, never used for training, never shared.
Guides

Start with your situation.

Each guide covers the format quirks, the specific tells and the practical next step for one scenario.

By generator

Checking audio from a specific tool?

The check itself is identical for all of them — our detector returns human-versus-synthetic and never names the generator. What changes is the context: where that tool's audio reaches people, and what the realistic risk is.

ElevenLabs OpenAI TTS Murf PlayHT Resemble AI
By role

What this looks like in your job.

Each of these has a different decision to make — and in every one of them, the control that matters most is not the detector.

FAQ

Common questions.

What is AI voice detection?
AI voice detection is the analysis of a recording to estimate whether the speech was produced by a human speaking into a microphone or generated by a text-to-speech or voice-cloning model. It works on the audio signal itself — spectral shape, micro-timing, breath and background noise — rather than on the words being said.
Is TextSight's voice detection free?
Yes. You get 3 audio checks per day with no signup at all. A free account raises that to 5 a day, Starter to 25 and Pro to 100. Business and Enterprise plans are unlimited.
What audio formats and file sizes are supported?
MP3, WAV, M4A, OGG and FLAC, up to 10 MB per file. The 10 MB limit is enforced on the server, so a larger file will be rejected regardless of plan. Opus files — the format WhatsApp uses for voice notes — are not accepted directly and need converting to MP3 or M4A first.
Can it tell me which AI tool generated the voice?
No. The check returns an AI-likelihood score and a human-versus-synthetic verdict. Engine attribution — naming ElevenLabs, OpenAI or another specific generator — is not something the detector returns today, and we do not display a guess we cannot stand behind.
Can I use the result as evidence in a legal or disciplinary case?
No. This is not forensic-grade analysis and it should never be the sole basis for a disciplinary, employment or legal decision. Treat the score as one input that tells you a recording is worth investigating, alongside provenance, context and corroboration from the people involved.
Does it work on languages other than English?
The analysis is based on acoustic characteristics of the audio rather than on the language being spoken, so it is not restricted to English. We have not published a separate measured accuracy figure per language, so treat non-English results with the same caution as English ones.
Is my audio stored?
No. Files are processed and discarded immediately. They are not stored, not used to train any model and not shared.
How long does a clip need to be?
A few seconds of clear, continuous speech is the practical minimum. Longer and cleaner clips give the detector more signal to work with. Very short clips, heavy background noise, aggressive compression and re-recording all reduce how much the result is worth.

Stop guessing whether a voice is real.

Three checks a day, free, no signup. Your audio is never stored.

Check an audio clip How accurate is it?
MP3 · WAV · M4A · OGG · FLAC · up to 10 MB · processed and discarded

Other TextSight detectors

Same approach, different medium.

AI Voice Detector — run a check AI Image Detector AI Text Detector AI Document Detector AI Hallucination Detector Limits of AI detection