Resemble AI is positioned at the enterprise end of synthetic speech, and it is unusual in selling both sides of the problem — voice cloning, and detection and watermarking tools alongside it. That combination is not a contradiction, and it points at something worth understanding: the long-term answer to synthetic audio is probably provenance rather than detection, and knowing the difference changes how much you should trust any score.
This is the part most pages about detecting Resemble AI voices quietly skip.
So what is this page for? Two things. First, the useful question is almost never “which tool made this” — it is “was this recording generated at all,” and that is the question a detector can actually answer. Second, knowing you are dealing with Resemble AI audio changes the context: where the recording is likely to have come from, what the realistic risk is, and what you should do next. That context is what the rest of this page covers.
If you see a tool that confidently attributes a clip to a named vendor, it is fair to ask what evidence that claim rests on, and how it performs against models released after it was trained.
Where its audio shows up, and why it also sells the countermeasure.
Resemble AI provides voice cloning and speech synthesis aimed at organisations rather than individual creators, and it has also built products on the other side of the problem — detection of synthetic speech and watermarking of generated audio. Several companies in this space have moved the same way.
The reason a synthesis vendor also builds watermarking is that blind detection has a structural ceiling, and everyone working on this knows it. A detector learns the traces left by the generators in its training data. New models leave different traces. Generation improves faster than detection corpora get rebuilt, so every deployed detector is permanently somewhat behind the newest systems. Care with your audio does not fix that; it is the honest limit of the whole category.
Provenance approaches invert the problem. Instead of examining audio for evidence of generation after the fact, they attach a signal at creation — an inaudible watermark, or signed metadata describing how a file was produced and edited. Industry efforts on content credentials and audio watermarking are both active, and several synthesis vendors now watermark their output by default.
| Blind detection | Provenance / watermarking | |
|---|---|---|
| Needs cooperation? | No — works on any file | Yes — only covers generators that participate |
| Unseen generators | Degrades, sometimes badly | Simply absent — no signal to find |
| Survives compression? | Poorly | Depends on the scheme; a design goal, not a given |
| Adversary who cares | Can post-process to evade | Can use a non-participating generator |
| Absence proves what? | Nothing conclusive | Nothing — no watermark is not evidence of authenticity |
Both have gaps, and the last row is the one people misread most often. Neither an absent watermark nor a “human” detector verdict establishes that a recording is genuine, unedited or in context. They are different partial signals, and the sensible position is to use whatever you have and to keep provenance — where the file came from and who handled it — as the primary evidence.
The check is the same for all of them — the context around it is what changes.
Why we publish no single accuracy percentage, what false positives look like, and how to weigh a result.
See our position →Run a known-genuine recording alongside the suspect one and read the gap. The strongest method available.
Learn the method →The technical guide: how detection works, and why it generalises poorly to generators it has not seen.
Read the technical guide →Three checks a day, free, no signup. Your audio is never stored.