A single score on a single file is the weakest way to use a voice detector, because it mixes up how the voice was produced with how the audio reached you. Running a recording you know is genuine through the same conditions gives you a baseline — and the gap between the two results is worth far more than either number on its own. This is the method, and it is free.
This is standard practice in any measurement discipline, and it is oddly rare in how people use AI detectors.
When you upload one clip and get one score, that score is responding to two independent things at once. The first is what you want to know: was this speech generated or physically recorded? The second is what you do not care about: how much of the original signal survived the phone codec, the messaging app, the forwarding, the re-encoding. You cannot separate them by looking harder at the number.
A control recording separates them for you. Take audio of the same person that you are confident is genuine, ideally arriving through a similar channel, and run it first. Whatever score it returns is your local definition of "real under these conditions". Now run the suspect clip. The distance between the two is the part that is actually about the voice.
It is not speaker verification. It does not tell you whether the two clips are the same person — our detector does not perform that task, and nothing on this page should be read as claiming otherwise. It tells you whether one clip carries more evidence of synthesis than another that you already trust.
It also does not detect editing. Two clips can both be entirely genuine audio, both score as human, and one of them can still be a misleading splice of real sentences. That is a provenance question, not a synthesis question.
The whole method rests on this choice. A bad reference is worse than no reference, because it produces a confident-looking comparison built on nothing.
The reference should have travelled a similar path to the suspect clip. Phone recording against phone recording. WhatsApp voice note against WhatsApp voice note. Studio file against studio file. This is what makes the comparison cancel out compression effects rather than simply measure them — a pristine WAV reference against a compressed voice note produces a large gap that is entirely about the codec and tells you nothing about the voice.
Where you cannot match the channel, say so to yourself before you interpret the numbers, and weight the result down accordingly.
Every difference you leave between the two files becomes a difference in the result that you will mistakenly attribute to the voice.
Check the reference first so you know what "real" looks like here before you see the suspect result. It is harder to read a number neutrally once you already have an answer in mind.
| Reference | Suspect clip | What it suggests |
|---|---|---|
| Clearly human | Clearly synthetic | The strongest result this method produces. The channel supports a confident reading and the suspect clip diverges sharply from a trusted baseline. Investigate properly — this is a reason to look, not a conclusion. |
| Clearly human | Clearly human | No synthesis evidence separating them. Note that this does not clear the recording of being edited, misattributed or taken out of context. |
| Uncertain or leaning synthetic | Similar to the reference | The audio conditions are driving both results. The comparison has not resolved the question. Try a better-quality pair or move to non-audio verification. |
| Uncertain | Clearly synthetic | Suggestive, but weaker than it looks. If the channel is degrading the reference this much, the suspect result deserves the same discount. Look for a cleaner reference before you lean on it. |
| Clearly synthetic | Anything | Stop. Either your reference is not genuine after all, or something about the pipeline is producing false positives. Re-examine the reference before reading anything into the comparison. |
Two reference clips from different occasions are noticeably better than one. If both baselines agree, you can trust the baseline. If they disagree with each other, the audio conditions are too variable for this comparison to mean anything, and you have learned something useful before drawing a wrong conclusion. Anonymous use allows three checks a day, which is exactly enough for two references and a suspect clip.
Reference: an ordinary voice note from the same person in the same chat, from weeks ago, about nothing important. Same app, same device, same compression — close to an ideal control. Convert both from Opus the same way, trim to similar lengths, run the old one first. If the old note reads clearly human and the new one clearly synthetic, that is a real finding. Then call the person back on a number you look up yourself anyway, because that settles it completely and this does not.
Reference: published, verifiable audio of the same speaker — a recorded interview or conference talk — ideally one that has been through a similar level of processing. The channel match is usually imperfect here, so treat the comparison as one strand among provenance, metadata and corroboration rather than as the deciding factor. The technical guide covers what else belongs in that assessment.
Reference: another recording of the same candidate from a different session on the same platform. Video-conferencing audio is heavily processed — noise suppression and automatic gain are doing a lot — so expect both scores to sit closer together than they would on raw audio, and be correspondingly slower to conclude anything. A live, unscripted conversation remains a better control than any recording comparison.
Where an old voice note from the same chat makes an almost perfect reference clip. Export and conversion steps.
Check a voice note →What moves reliability up and down, what false positives look like, and how much a gap is worth.
See our position →Why channel mismatch dominates results, and what else belongs in a serious assessment of a recording.
Read the technical guide →Three checks a day, free, no signup. Your audio is never stored.