Short answer: it depends so heavily on the audio that a single percentage would tell you almost nothing, which is why we do not publish one. Long answer below — what actually moves reliability up and down, what a false positive looks like and who it lands on, and how to weigh a result before you act on it. This is the page we would want to read before trusting anyone else's detector.
This is a deliberate choice, and it costs us conversions against competitors who do. Here is the reasoning.
An accuracy figure is a measurement taken under conditions. Change the conditions and the figure changes — not slightly, but by margins large enough to reverse the conclusion you would draw from it. In voice detection the conditions that matter are which generator produced the audio, how the audio reached you, how long it is, and how much processing sits between the microphone and your file. None of those are constant, and none of them are things we know about your clip.
So "99% accurate" answers a question nobody is asking. The question you actually have is "how much should I trust this result, on this file" — and a headline number cannot answer it. Worse, it actively interferes, because it invites you to apply a laboratory figure to a forwarded voice note where it does not remotely hold.
These apply to every detector in this category, ours included. Use them to judge your own clip before you judge the answer.
The last item on the right deserves emphasis, because it is structural rather than incidental. Detection models learn the traces left by the synthesis systems they were trained against. Generative speech models are released faster than detection models are retrained. Every deployed detector is therefore permanently somewhat behind the newest generators, and no amount of care with your audio fixes that. It is the honest ceiling on this whole category.
A missed deepfake is a failure. A wrongly flagged real person is a different kind of failure, and it is the one worth designing against.
A false positive is a genuine human recording scored as synthetic. The asymmetry is the point: when a deepfake slips through, you are back where you started. When a real recording is flagged, someone now has to prove a negative about their own voice — to an employer, a platform, a client, or a family member who has already decided.
| Situation | Why it reads as synthetic |
|---|---|
| Professionally produced audio | Noise reduction, compression, de-essing and EQ remove exactly the untidiness that marks a recording as physical. A polished podcast can look cleaner than reality allows. |
| Treated rooms and close micing | A voice booth has almost no room tone. The absent noise floor is a genuine property of a good recording space and also a synthesis tell. |
| Very short clips | Not enough speech to establish a pattern, so the score drifts toward whichever side the audio conditions happen to favour. |
| Even, measured delivery | Someone reading prepared remarks calmly and clearly produces regular prosody. Regularity is one of the things the detector is looking for. |
| Heavily transcoded uploads | Audio that has been through several platforms carries stacked codec artefacts that a model has no principled way to distinguish from generation artefacts. |
Notice the pattern: several of these describe people who record well. Voice actors, podcasters, broadcasters, anyone using a decent microphone in a treated space — professionals are more exposed to false positives than someone talking into a phone in a kitchen. That is a real fairness problem in this category and it deserves saying out loud rather than being buried.
| Result | Reasonable reading | Not a reasonable reading |
|---|---|---|
| Confidently synthetic, clean audio | Strong reason to treat the recording as untrusted and to investigate its provenance seriously. | Proof that a named person faked it, or grounds on its own for a disciplinary or legal action. |
| Confidently synthetic, compressed audio | Worth acting on cautiously — compression usually pushes scores toward uncertainty, so a confident result despite it is meaningful. | Equivalent to the clean-audio case. The evidence base is thinner. |
| Confidently human | The audio carries the physical signature of a real capture. | That the recording is honest, unedited, in context, or of the person you think it is. Detection answers none of those. |
| Uncertain | The file did not carry enough usable signal. Almost always a statement about the audio, not about the voice. | Weak support for whatever you already believed. This is the most common misreading and the most costly. |
Concretely, the things it should not decide by itself:
In each case the detector's proper role is to tell you where to spend your attention. Where the stakes genuinely justify it, a qualified forensic audio examiner can do what an automated score cannot: examine chain of custody, container structure and edit points, and defend a stated methodology with stated uncertainty.
We do not name the generator. Attributing audio to ElevenLabs, OpenAI or any specific vendor is a much harder problem than telling human from synthetic, and our detector does not return it. We would rather show nothing than a plausible guess.
We do not claim per-language parity. Detection works on acoustics rather than language, so it is not English-only, but we have not published measured figures per language and will not imply an equivalence we have not tested. The Hindi guide says the same thing in Hindi.
The technical version of this page: EER, cross-domain evaluation, and why detectors generalise poorly to unseen generators.
Read the technical guide →The control-recording method, and why it beats an isolated score on almost every real-world file.
Learn the method →The same honesty applied to text detection — false positives, thresholds and why no score is proof.
Read the explainer →Three checks a day, free, no signup. Your audio is never stored.