HomeVoice Detection › Compare Two Clips

Compare two audio clips — is one of them AI?

A single score on a single file is the weakest way to use a voice detector, because it mixes up how the voice was produced with how the audio reached you. Running a recording you know is genuine through the same conditions gives you a baseline — and the gap between the two results is worth far more than either number on its own. This is the method, and it is free.

Open the voice detector Pick a reference clip ↓
Two checks is enough 3 free checks a day
The idea

Why a control clip changes everything.

This is standard practice in any measurement discipline, and it is oddly rare in how people use AI detectors.

When you upload one clip and get one score, that score is responding to two independent things at once. The first is what you want to know: was this speech generated or physically recorded? The second is what you do not care about: how much of the original signal survived the phone codec, the messaging app, the forwarding, the re-encoding. You cannot separate them by looking harder at the number.

A control recording separates them for you. Take audio of the same person that you are confident is genuine, ideally arriving through a similar channel, and run it first. Whatever score it returns is your local definition of "real under these conditions". Now run the suspect clip. The distance between the two is the part that is actually about the voice.

Concretely: if a known-real voice note from your brother comes back leaning synthetic, that tells you WhatsApp compression is pushing scores in that direction for this speaker on this channel — and a similar-looking result on the suspicious note means much less than it appeared to. If the reference comes back clearly human and the suspect clip clearly synthetic, you have something worth acting on.

What the comparison is not

It is not speaker verification. It does not tell you whether the two clips are the same person — our detector does not perform that task, and nothing on this page should be read as claiming otherwise. It tells you whether one clip carries more evidence of synthesis than another that you already trust.

It also does not detect editing. Two clips can both be entirely genuine audio, both score as human, and one of them can still be a misleading splice of real sentences. That is a provenance question, not a synthesis question.

Step 1

Choosing a reference clip you can trust.

The whole method rests on this choice. A bad reference is worse than no reference, because it produces a confident-looking comparison built on nothing.

Good reference

Independently trustworthy

  • A video call you personally attended and recorded
  • A voicemail from well before the incident
  • A published interview, podcast or conference talk
  • A voice note from an unrelated, ordinary conversation
  • Anything whose authenticity does not depend on the question you are investigating
Not a reference

Circular or mismatched

  • Another clip from the same suspicious source
  • Anything the person under question supplied to prove themselves
  • A clip from the same conversation as the suspect audio
  • Studio-produced audio compared against a phone recording
  • A different speaker entirely
The circularity trap. If both clips came from the same source, matching scores prove nothing. Two synthetic clips from the same generator will score similarly, and you will read that similarity as reassurance. The reference must come from somewhere the suspect clip could not have come from.

Match the channel, not just the speaker

The reference should have travelled a similar path to the suspect clip. Phone recording against phone recording. WhatsApp voice note against WhatsApp voice note. Studio file against studio file. This is what makes the comparison cancel out compression effects rather than simply measure them — a pristine WAV reference against a compressed voice note produces a large gap that is entirely about the codec and tells you nothing about the voice.

Where you cannot match the channel, say so to yourself before you interpret the numbers, and weight the result down accordingly.

Step 2

Prepare both clips identically.

Every difference you leave between the two files becomes a difference in the result that you will mistakenly attribute to the voice.

  1. Take a similar length from each. Aim for thirty seconds or more of continuous speech from both. If one clip is only eight seconds, trim the other to roughly match rather than giving the detector far more to work with on one side.
  2. One speaker only. Cut out the other side of a conversation, background music and long silences from both files.
  3. Export both in the same format at the same bitrate. Same container, same encoder settings. This is the step people skip, and it is the one that most often manufactures a fake gap.
  4. Do not process one and not the other. No noise reduction, normalisation or EQ on either — and certainly not on just one. Any cleanup you apply to a single file is a change you will read as a finding.
  5. Keep the originals. Work on duplicates. If this ever becomes a real dispute, the untouched files are what matter.
Format reminder: the detector accepts MP3, WAV, M4A, OGG and FLAC, up to 10 MB per file. WhatsApp voice notes arrive as Opus and need converting first — the WhatsApp guide covers that. Convert both clips the same way.
Step 3

Run both, then read the gap.

Check the reference first so you know what "real" looks like here before you see the suspect result. It is harder to read a number neutrally once you already have an answer in mind.

ReferenceSuspect clipWhat it suggests
Clearly human Clearly synthetic The strongest result this method produces. The channel supports a confident reading and the suspect clip diverges sharply from a trusted baseline. Investigate properly — this is a reason to look, not a conclusion.
Clearly human Clearly human No synthesis evidence separating them. Note that this does not clear the recording of being edited, misattributed or taken out of context.
Uncertain or leaning synthetic Similar to the reference The audio conditions are driving both results. The comparison has not resolved the question. Try a better-quality pair or move to non-audio verification.
Uncertain Clearly synthetic Suggestive, but weaker than it looks. If the channel is degrading the reference this much, the suspect result deserves the same discount. Look for a cleaner reference before you lean on it.
Clearly synthetic Anything Stop. Either your reference is not genuine after all, or something about the pipeline is producing false positives. Re-examine the reference before reading anything into the comparison.

Add a second reference when it matters

Two reference clips from different occasions are noticeably better than one. If both baselines agree, you can trust the baseline. If they disagree with each other, the audio conditions are too variable for this comparison to mean anything, and you have learned something useful before drawing a wrong conclusion. Anonymous use allows three checks a day, which is exactly enough for two references and a suspect clip.

Still not proof. A large gap is stronger evidence than an isolated score — that is the entire point of doing this — but it remains a triage signal, not forensic-grade analysis, and it must never be the sole basis for a disciplinary, employment, financial or legal decision. Our full position on what a result is worth is here.
In practice

Three situations, worked through.

A voice note asking for money

Reference: an ordinary voice note from the same person in the same chat, from weeks ago, about nothing important. Same app, same device, same compression — close to an ideal control. Convert both from Opus the same way, trim to similar lengths, run the old one first. If the old note reads clearly human and the new one clearly synthetic, that is a real finding. Then call the person back on a number you look up yourself anyway, because that settles it completely and this does not.

A leaked clip before publication

Reference: published, verifiable audio of the same speaker — a recorded interview or conference talk — ideally one that has been through a similar level of processing. The channel match is usually imperfect here, so treat the comparison as one strand among provenance, metadata and corroboration rather than as the deciding factor. The technical guide covers what else belongs in that assessment.

A recorded screening call

Reference: another recording of the same candidate from a different session on the same platform. Video-conferencing audio is heavily processed — noise suppression and automatic gain are doing a lot — so expect both scores to sit closer together than they would on raw audio, and be correspondingly slower to conclude anything. A live, unscripted conversation remains a better control than any recording comparison.

FAQ

Comparison questions.

Why compare two clips instead of just checking one?
Because an isolated score mixes together two things you cannot separate: how the voice was produced, and how the audio reached you. Running a known-genuine recording through the same channel conditions gives you a baseline for what "real" looks like in this specific situation, so the difference between the two scores isolates the part you actually care about.
Does the detector compare two files for me?
No. The tool analyses one clip at a time and returns a score for it. The comparison is something you do by running two checks and reading the results side by side. On the free tier you get 3 checks a day, which covers a reference and a suspect clip with one to spare.
Does this tell me whether both clips are the same person?
No. That is speaker verification, a different task, and our detector does not perform it. The comparison method tells you whether one clip carries more synthesis evidence than another — not whether the two voices belong to the same individual.
What makes a good reference clip?
One whose authenticity does not depend on what you are investigating, captured through a similar channel to the suspect clip, of similar length, with the same person speaking naturally. A recorded video call you personally attended is close to ideal. Anything the suspect party supplied to you is not a reference at all.
What if both clips score the same?
Then the audio conditions are dominating the result and the comparison has not resolved anything. Try a longer or less-compressed pair, or accept that the question cannot be answered from this audio and move to non-audio verification instead.
Can I use a clip the other person sent me as the reference?
No, if that person is who you are investigating. A reference has to be independently trustworthy. If both clips came from the same source, a matching score tells you only that they were produced the same way — which is exactly what you would expect if both were synthetic.
How many checks does this take?
Two at minimum — one reference and one suspect clip. Three if you want a second reference from a different occasion, which is worth it when the stakes are high. Three checks a day, free, whether or not you are signed in. Voice has its own plans, separate from our text subscriptions — Echo $9/mo for 100 checks a month up to Broadcast at $99 for 10,000 — plus one-time Soundbites credit packs from $5.
Is a large gap between scores proof?
No. It is stronger evidence than a single score, which is the point of the method, but it is still not forensic-grade and should never be the sole basis for a disciplinary, employment, financial or legal decision. It tells you the suspect clip differs from a known-genuine baseline in a way worth investigating properly.
Related

More voice detection guides.

Run the reference first. Then the suspect clip.

Three checks a day, free, no signup. Your audio is never stored.

Open the voice detector All voice guides
One clip at a time · compare the results yourself · a gap is evidence, never proof