For a fraud or security function the useful question is not “can we detect deepfakes” but “where does our process still treat a voice as proof of identity?” Voice has stopped being an authentication factor, and the controls that follow from accepting that are procedural. Detection has a real but narrow role, mostly after the fact.
Everything sensible follows from accepting this, and nothing sensible follows from resisting it.
A recognisable voice used to be reasonable evidence that a known person was on the line. It is not any more, and no detection capability restores it — the arms race structurally favours generation, because synthesis improves faster than detection corpora are rebuilt.
So the design question is not how to spot cloned voices. It is where your processes currently allow a voice, on its own, to authorise something. Common places that survive longer than they should:
None of these depend on detecting anything.
Detection earns its place after the event rather than during it.
Where a call was recorded, checking the audio adds a line to the incident record and helps characterise what happened. That matters for internal reporting, for insurance and for law enforcement referral — as supporting context, explicitly not as a determination.
Across a set of related incidents, consistent findings help establish whether you are seeing one campaign or several. Individually weak signals can be collectively informative.
Genuine examples from your own environment are considerably more effective in awareness training than generic ones, and a check helps you label them correctly before you use them.
What to do in the first sixty seconds, the tells of a cloned voice, and where to report it.
What to do now →Why we publish no single accuracy percentage, what false positives look like, and how to weigh a result.
See our position →What the public API covers today, what it does not, and how to get in touch about volume.
See the API status →Three checks a day, free, no signup. Your audio is never stored.