Cloned voices are now good enough to fool the ear, and a few seconds of someone speaking is enough raw material to build one. This is the practical guide: how detection actually works, what a score is worth, what it is not worth, and how to handle the specific situations people land here with — a WhatsApp voice note, a scam call, a Hindi recording, a clip you need to verify before you publish it.
It reads the recording, not the words. A detector never considers what was said or whether the claim is plausible — it looks at how the sound was produced.
A human voice reaches a file through a long, messy physical chain. Air moves through a throat and mouth, bounces around a room, hits a microphone diaphragm, and passes through a preamp and a codec. Every stage adds irregularity: room reflections, breath, lip and tongue noise, tiny timing drift, a noise floor that never quite goes silent.
Synthetic speech skips almost all of that. A text-to-speech or voice-cloning model reconstructs a waveform directly from a learned statistical model of what speech looks like. It is very good at reproducing the parts of a voice we consciously notice — timbre, accent, intonation — and much weaker at reproducing the parts we do not, because those parts were never the point of the training objective.
Detection lives in that gap. The result is a probability, not a fact, and it is a probability about how the audio was produced — never about whether the person in the recording is trustworthy or whether what they said is true.
| Signal | What a real recording looks like | What often gives synthesis away |
|---|---|---|
| Spectral detail | Energy scattered untidily across the frequency range, including high frequencies the codec barely preserves | Smoother, more regular spectral shape; sometimes a hard ceiling where the vocoder stops generating detail |
| Micro-timing | Pauses between words vary constantly, even mid-sentence | Gaps that land in a narrower, more uniform range than a person can produce |
| Breath and mouth noise | Audible inhales, lip smacks, swallows, the odd stumble | Breath either absent or placed identically each time it appears |
| Noise floor | Continuous room tone that persists through silences | Silences that are cleaner than the room ever was, or noise that starts and stops with speech |
| Energy contour | Volume rises and falls unevenly inside each syllable | Envelopes that repeat a learned shape across different words |
No single one of these is decisive. A close-miced studio recording in a treated room can look suspiciously clean; a synthetic clip played through a phone speaker and re-recorded picks up genuine room noise that masks the tells. The verdict comes from weighing them together, which is exactly why it arrives as a confidence score rather than a yes or a no.
Most of the harm in this category comes from people treating a detector output as proof. It is not proof, and every honest vendor in the space will tell you the same thing.
These apply to every detector on the market, ours included. If two or more are true of your clip, weight the result much less heavily:
You will see competitors advertise numbers like "99% accurate". Treat those with suspicion unless they name the benchmark, the generators tested, the audio conditions and the false-positive rate. Detection accuracy is not one number — it varies enormously with how the audio reached you and whether the generating model resembles anything the detector was trained against. We explain our position on this in full, including why a number without conditions attached is worse than no number.
Naming a specific generator — ElevenLabs, OpenAI, Murf — from audio alone is a much harder problem than telling human from synthetic, and our detector does not return it. Rather than display a plausible-sounding guess, we leave it out. If you see a tool confidently attributing a clip to a named vendor, ask what it is basing that on.
Each guide covers the format quirks, the specific tells and the practical next step for one scenario.
How to export a voice note on iOS and Android, why the Opus format needs converting first, and what compression does to the result.
Check a voice note →What to do in the first sixty seconds, the tells of a cloned-voice call, and how to report it in India, the US and the UK.
What to do now →क्या यह आवाज़ AI से बनी है? How language affects detection, and what we can and cannot claim about Hindi performance.
हिंदी गाइड पढ़ें →The technical view: how audio deepfakes are built, why detectors generalise poorly to unseen generators, and what the research actually shows.
Read the technical guide →The reference-clip method: run a known-real recording of the same person alongside the suspect one and read the gap between them.
Learn the method →Why we do not publish a single accuracy percentage, what false positives look like, and how to weigh a result honestly.
See our position →The check itself is identical for all of them — our detector returns human-versus-synthetic and never names the generator. What changes is the context: where that tool's audio reaches people, and what the realistic risk is.
Each of these has a different decision to make — and in every one of them, the control that matters most is not the detector.
Where detection fits in a verification workflow, what it cannot establish, and how to describe a result in copy so the hedge survives editing.
The newsroom workflow →Voice has stopped being an authentication factor. The process controls that follow from accepting that, and detection's narrow retrospective role.
Controls that work →Why a detector is the weakest part of your defence against candidate impersonation, and what to do instead without wrongly accusing a real applicant.
Hiring controls →Disclosure rules, what to do if your own voice is cloned, and why professionally produced audio is more likely to be wrongly flagged than a phone recording.
Creator guide →Five questions that separate a real capability from a marketing claim, the answers that should worry you, and a ten-minute test you can run yourself.
Buyer's guide →Straight answer on what the public API covers today, what it does not, and how to reach us with a volume use case.
API status →Three checks a day, free, no signup. Your audio is never stored.
Same approach, different medium.