A clip lands in your inbox. Forty seconds, a recognisable voice, saying something that would be a story. The sender will not say where it came from. Your editor wants to know by end of day whether it is real.
The instinct is to run it through a detector first. I would put that fifth, and I run a detection product.
The reasoning is straightforward. A detector answers one narrow question, and it answers it with the least reliable evidence in your possession. Everything else you can establish about that file is more robust, more explainable to a lawyer, and more useful to a reader. So work the strong evidence first and let the audio analysis inform a picture you have already built.
Here is the order I would work in.
1. Establish what you actually have
Before anything else, freeze the artifact.
Get the original file, not a re-share. If it arrived through a messaging app, ask for it as a document or file attachment rather than as a voice message, because most platforms re-encode audio sent through the voice pathway. If the person has an earlier copy, ask for that instead. Every hop between the original recording and your copy destroys evidence, and you cannot get it back later.
Hash the file immediately, note the time you received it, and store it read-only. If this becomes a legal matter, the fact that you can demonstrate the file has not changed since you got it will matter more than any analysis you run on it.
Then write down what you know about the path it took to reach you. Not what you were told. What you can verify. In most cases this is a short and unsatisfying list, and writing it down is what stops you from quietly upgrading a claim to a fact three days later.
2. Read the container before you listen to the content
Audio files carry metadata, and it is worth looking before you do anything else.
Depending on format you may find an encoder string, a creation timestamp, a sample rate and bit depth, channel configuration, and sometimes device or application identifiers. Command line tools like ffprobe or exiftool will show you all of it in a few seconds.
What to do with it:
An encoder string that names a specific recording application, and a sample rate consistent with that application's defaults, is weak corroboration of the story you were told. An encoder string that names an audio editor is a flag worth pulling. It does not mean the content was fabricated, since perfectly innocent conversions pass through editors, but it means the file you are holding is not a camera-original and you should say so internally.
A duration that does not match a natural start and stop, or a file whose bitrate profile changes partway through, can indicate assembly. Splices are frequently audible too, if you know to listen for the background ambience cutting rather than continuing.
The critical rule: missing metadata proves nothing. Most platforms strip it as a matter of course. An absence is the expected state, not a finding. I have watched people build a theory on stripped metadata and it is always embarrassing.
3. Look for content provenance, and understand what its absence means
There is now a real technical standard for cryptographically signed provenance, developed by the C2PA coalition (c2pa.org). Where it is present, it can tell you what device or application produced a file and what has been done to it since, with a signature you can check.
If a clip carries valid C2PA provenance, that is genuinely strong evidence and much better than anything a detector will give you. Check it.
The problem is coverage. Adoption is partial, most audio in circulation carries nothing, and provenance data does not survive most re-sharing. So in practice you will usually find nothing, and finding nothing tells you nothing at all. Provenance is a strong positive signal and a meaningless negative one. Do not let anyone in the newsroom treat "no provenance data" as suspicious.
4. Corroborate against the world
This is the part that actually decides most stories, and it has nothing to do with audio analysis.
Does the content check out? If the speaker references an event, a document, a meeting, a decision, can you independently establish those happened and happened in the order described? Fabricated audio is usually generated from a script written by someone with incomplete knowledge, and incomplete knowledge shows up as small factual errors about things the real speaker would have known precisely.
Does the acoustic environment match the claimed setting? Room reverberation is hard to fake convincingly and easy to reason about. A clip that is supposed to be a phone call from a car should have road noise, an appropriate frequency profile, and reverb consistent with a small enclosed space. A clip that is supposed to be a boardroom should not sound like a close-miked studio. Background details are frequently where fabrications fall apart, because the fabricator was thinking about the words.
Who else should have heard this? If the recording captures a meeting, other people were in that meeting. Any of them can confirm or deny that the exchange happened, without you ever having to prove anything about the file. This is often the fastest route to an answer and it is routinely skipped in favour of technical analysis, which is backwards.
Where else has this circulated? Search for the clip, for quotes from it, for reporting on it. Prior circulation with a different attributed context is a very strong signal, and recycled audio presented in a new frame is more common than outright synthesis.
5. Now run the detector
With the picture built, audio analysis becomes useful, because you know what you are testing.
Our detector takes MP3, WAV, M4A, OGG and FLAC files up to 25 MB. It returns an AI score from 0 to 100, a human score that is exactly its complement, and a verdict of human, AI, or uncertain. You can run a clip free without an account at the voice detector.
Three things a newsroom needs to understand about what comes back.
Uncertain is a real answer, and you will see it. Before scoring, the detector checks whether the clip carries enough usable speech. Too short, too quiet, or not enough actual speech and it declines to judge rather than producing a number. Leaked audio is very often exactly this material: brief, noisy, heavily compressed. When you get uncertain, the finding is that this file cannot support a technical claim in either direction. That is publishable as a limitation and it is much better than a fabricated confidence.
A low AI score is not proof the audio is genuine. It means no synthesis artifacts were found. Artifacts are destroyed by compression, by re-encoding, by noise, and by post-processing, all of which describe the average leaked clip. Absence of evidence, in this domain particularly, is weak evidence of absence. Meanwhile a clip can be entirely genuine audio that has been deceptively edited, and a detector will happily score real speech as real speech no matter what order the sentences are in.
It will not name the tool. We do not report which product generated a clip, because that claim cannot be recovered from the waveform. The artifacts belong to vocoder architectures, and architectures are shared and swapped between products constantly. I set out the full argument in the engine comparison piece. If a tool tells you a clip was made with a specific named product, ask how it distinguishes that product from every other one running similar architecture.
Run more than one detector if you can. Where they agree you have a stronger signal. Where they disagree, that disagreement is itself the finding, and why voice detectors disagree explains the mechanics well enough to quote in a methods note.
6. Consider the boring explanations first
Most disputed audio is not synthetic. Ranked roughly by how often I would expect to see them:
Real audio, deceptively edited. Sentences reordered, a qualifier cut, an answer attached to a different question. No detector catches this because there is nothing synthetic in it.
Real audio, correct content, wrong context. Genuinely said, years ago, about something else.
Real audio of an impersonator. A human doing an impression, which some detectors handle badly and which is far cheaper than cloning.
Real audio, misidentified speaker. Two people with similar voices and a confident assumption.
Synthetic audio. Genuinely happens, genuinely worth checking for, and still not the first hypothesis on the list.
If your verification process only looks for the last one, it will miss the four above it.
Questions worth asking the source
Source handling is verification work, and in my experience it produces more usable evidence than any technical step on this list. A few questions that tend to be productive.
How did this reach you? Not who made it. How it arrived. Every hop between origin and your source is a hop you can potentially trace, and each one is a person who might talk.
Is this the file you received, or a copy you made? Sources routinely re-export, trim, or convert audio before passing it on, often to protect themselves, and they usually do not think to mention it. Asking directly gets you a better file surprisingly often.
Do you have anything from before this? A longer version. The message it was attached to. A screenshot of the conversation where it appeared. Context around a file is frequently easier to verify than the file itself, and harder to fabricate convincingly.
Who else has this? If a clip has circulated in a group, other recipients can independently confirm when it appeared, which establishes a floor on its age. That alone can kill a fabrication theory or a genuine one.
What do you think it is? Open ended on purpose. A source who has already convinced themselves of an interpretation will tell you, and knowing their theory helps you notice where you are being led.
None of this requires trusting the source. It requires treating their account as a set of claims that generate checkable predictions, which is the same posture you would take with any other evidence.
What to write in the piece
If you publish, describe your verification in terms a reader can evaluate.
Say where the file came from and what you could and could not establish about that path. Say what you checked. If you ran detection, name the limitation rather than the score alone: something like an AI voice detection check returned an uncertain result, which is expected for audio of this length and compression level, is honest and informative. A bare percentage in a news story implies a precision the method does not have.
If you could not verify, and the story runs anyway because the content is newsworthy regardless, say that plainly and high up. Readers handle acknowledged uncertainty better than they handle discovering later that it was hidden.
The failure mode to avoid
I will end with the thing I actually worry about, which is not fabricated audio getting published.
It is the reverse. As deepfakes become a familiar concept, genuine recordings get dismissed as fakes, and the dismissal is now cheap and socially acceptable. A subject of a story can simply say the audio was AI generated and shift the burden onto a newsroom that may not be able to prove otherwise, especially given that a low AI score is not proof of authenticity and compression can strip whatever evidence existed.
The defence against that is not a better detector. It is documentation. A clear, contemporaneous record of where a file came from, who handled it, what was checked, what was found, and what remained unknown. That record is what stands up when someone tries to make the question unanswerable.
Build it while you are working, not afterwards. It takes an extra ten minutes and it is the difference between a verification you can defend and a score you cannot.
Try it on your own writing