There is a version of this article that opens by telling you how frightening voice cloning has become, quotes a scary statistic, and then sells you a detector. I want to write the other version, because after looking at how the documented cases actually unfolded, I do not think detection is where most of the defence belongs.
Two incidents are worth knowing in detail. Not because they are the only ones, but because both are properly sourced rather than vendor folklore.
The two cases people actually cite
In March 2019, the chief executive of a UK energy firm took a call he believed was from the head of the German parent company. The voice had the right accent and, in his own description, the right melody. He was told to move 220,000 euros to a Hungarian supplier within the hour. He did. The money moved onward to Mexico and then elsewhere. The story reached the public through the company's insurer, Euler Hermes, and was reported by the Wall Street Journal and picked up widely afterwards (Forbes, Sophos).
The detail I keep coming back to is what stopped it. Not the first transfer. The attackers came back for a second and a third, and it was the repetition that broke the spell. The audio never got worse. The pattern got suspicious.
In early 2024, an employee at the engineering firm Arup, in Hong Kong, joined a video call with people who looked and sounded like the company's UK-based CFO and several familiar colleagues. Every participant was fabricated. He authorised transfers totalling roughly 200 million Hong Kong dollars, about 25.6 million US dollars (CNN, February 2024; CNN, May 2024, naming Arup).
Read the reporting on the Arup case closely and one thing jumps out. The employee had already been suspicious. He received a written message about a confidential transaction and thought it looked like phishing. Then he joined the call, saw and heard people he recognised, and set the suspicion aside.
The technology did not defeat a control. It defeated a doubt that a human being had already correctly formed.
Why detection is the wrong first line
I run a voice detection product, so treat this as an argument against my own commercial interest.
Our detector, and every detector I know of, works on a file. You upload audio, it analyses it, it returns a score. That model does not fit the moment of the attack. Nobody on a live call is going to say hold on, let me record this, export it, and upload it to a website before I answer you. And in the Arup case there was no suspicious file to analyse at all, because the interaction was a live video conference.
There is a second problem, which is that phone and conference audio is exactly the hard case for detection. Traditional telephony discards everything above roughly 3.4 kHz, and conferencing platforms apply aggressive codecs, noise suppression and packet loss concealment. Most of the fine spectral evidence a detector relies on has been processed away before you could analyse it. I go into the specifics of this in our piece on cloned voices over phone calls, and the honest summary is that live call audio is the worst material you can hand a detector.
So detection has a role here, but it is a forensic role after the fact, or a triage role on a voicemail or a forwarded recording. It is not the thing standing between your company and a wire transfer.
The thing standing between your company and a wire transfer is your payment process. And in both documented cases, the process is what failed.
The control that would have stopped both
If you take one thing from this article, take this: identity should not authorise payments. Procedure should.
The attack in both cases works by convincing a person that they know who is speaking. Every control that depends on recognising a voice, a face, a manner, or a relationship is a control that voice cloning is specifically built to defeat. You cannot train your way out of it, because the whole point is that the imitation is good.
What cloning cannot defeat is a rule that says: this class of payment requires these steps, and the steps do not change based on who is asking.
Here is what that looks like concretely.
Out-of-band callback to a stored number. Any payment above your threshold, or any change to bank details, requires a verification call placed by your team to a number already on file from before the request. Not a number in the email. Not a number the caller gives you. Not a callback to the number that just called you. The direction of the call matters enormously: attackers control inbound, they do not control where you dial.
Dual authorisation with genuine independence. Two approvers, and the second one must be able to say no without professional consequence. In practice this is where the control quietly dies. If your second approver is junior to the person requesting, and the request is framed as urgent and confidential and coming from senior leadership, you do not have dual authorisation. You have one approver and a witness.
A cooling period on new payees. New bank details do not become payable for a fixed window, commonly 24 or 48 hours. This is unglamorous and it is probably the highest value item on the list. Fraud proceeds are moved fast because they have to be. A delay that is trivial for legitimate business is fatal to the attack, and it also creates the window in which the second and third requests, the ones that broke the 2019 case, become visible as a pattern.
Confidentiality as a red flag, not a reason. Both documented cases used secrecy to prevent the target from checking. Write it into policy: a request to bypass normal approvals because a transaction is confidential is itself grounds to escalate. Say it out loud, repeatedly, to everyone in a position to move money. The fraud depends on the target believing that checking would be insubordinate.
A named escalation path with no penalty. Everyone who can initiate a payment should know exactly who to call when something feels wrong, and should have heard, from leadership, that pausing a legitimate transaction has never once cost anyone here anything. This is a cultural control and it is worth more than a lot of software.
Where a shared phrase helps, and where it does not
Some teams adopt a spoken code word for verifying identity on calls. It is a reasonable low-cost measure and it does something real, but be clear about its limits.
It works against an attacker who has cloned a voice from public material and knows nothing about your internal conventions. It does not work against an attacker who has compromised a mailbox or a chat account and read your last six months of messages, which is a very common precondition for this kind of fraud. Voice fraud usually sits on top of an existing intrusion rather than arriving out of nowhere.
Treat a shared phrase as a speed bump, not a gate. The callback to a stored number is the gate.
What to do when a call feels wrong, in the moment
The tactical advice, for the person actually on the call:
End the call yourself and dial back on a known number. This alone defeats most of the attack surface. If the person on the other end objects to being called back on their own known number, you have your answer.
Ask something contextual that is not in any document. Not a security question, just a normal human reference to something specific and recent that would not appear in an email thread. Voice models generate speech, they do not supply the speaker's memory, and an attacker working from a script has to improvise.
Notice what is being asked of you, not just who is asking. Urgency, confidentiality, and a bypass of normal process appearing together is the signature. That combination is worth more as a signal than anything in the audio.
If you can, record. On many systems a call can be recorded or a voicemail preserved. This is where detection becomes useful, and it is also the evidence your bank and the police will want. Which brings me to the part where the tool does earn its place.
Where audio detection genuinely helps
After the fact, and on files rather than live calls.
If you have a preserved voicemail, a forwarded voice note, or a recording of the call, running it through a detector gives you an additional signal for the incident report. Our detector accepts MP3, WAV, M4A, OGG and FLAC files up to 25 MB, and returns an AI score from 0 to 100 with a matching human score and a verdict of human, AI, or uncertain.
That uncertain verdict is going to come up more often than you would like on this material, and I would rather explain why than have you think the tool is broken. Phone recordings are short, narrowband, and heavily compressed. Our detector runs a quality check before scoring and declines to judge clips that do not carry enough usable speech. When it says uncertain, it means the evidence in the file cannot support a claim in either direction, which is a true and useful thing to know even though it is not the answer you wanted.
The detector also will not tell you which product generated the voice. That claim is not recoverable from the audio, for reasons I set out in the engine comparison piece. Attribution in a fraud case comes from payment trails and platform records, not from a waveform.
Use it as one line in the incident file, alongside the transfer records and the message headers. Not as the finding.
A twenty minute exercise worth doing this quarter
Get your finance team in a room and walk a scenario, out loud, with no slides.
The CFO calls the AP clerk at 4:40pm on a Friday. The voice is right. There is a confidential acquisition, the deal collapses if it leaks, and a deposit has to reach a new account before end of day. The CFO is boarding a flight and cannot take questions.
Now ask, specifically: what happens next in your organisation, procedurally, not aspirationally? Who does the clerk call? Do they have that number without asking someone? Is the second approver senior enough to refuse? Would the payment actually be blocked by a new-payee hold, or is there an override, and who can use it?
Most teams find at least one gap in twenty minutes. In my experience the gap is usually the same one: the escalation path exists on paper but the person who would need to use it has never used it, does not have the number to hand, and would feel awkward being the one to slow down something the CFO asked for at 4:40 on a Friday.
That awkwardness is the actual vulnerability. Not the audio.
The uncomfortable summary
Voice cloning has made a category of social engineering cheaper and more convincing, and it will keep getting better. The defence that scales is not better ears, human or machine. It is a payment process where knowing who is on the phone does not, by itself, get money out the door.
Both documented cases involved a person who had good instincts. One of them acted on the instinct at transfer two and three. The other set it aside when the faces looked familiar. Build the process so that neither of those outcomes depends on an individual's confidence in a moment of manufactured pressure.
If you have a recording you want a second opinion on, our voice detector is free to try without an account. Just do not let it be the last step before the money moves. It should be the step after.
Try it on your own writing