HomeVoice Detection › Resemble AI

Detect Resemble AI voices.

Resemble AI is positioned at the enterprise end of synthetic speech, and it is unusual in selling both sides of the problem — voice cloning, and detection and watermarking tools alongside it. That combination is not a contradiction, and it points at something worth understanding: the long-term answer to synthetic audio is probably provenance rather than detection, and knowing the difference changes how much you should trust any score.

Check an audio clip free Provenance vs detection ↓
3 checks a day, no signup Audio never stored
Read this first

What we can and cannot tell you.

This is the part most pages about detecting Resemble AI voices quietly skip.

Our detector does not name the generator. It returns an AI-likelihood score and a human-versus-synthetic verdict. It will not tell you “this was Resemble AI” — that is source attribution, a substantially harder problem than telling synthetic from real, and our detector does not return it. We would rather show you nothing than a plausible-sounding guess.

So what is this page for? Two things. First, the useful question is almost never “which tool made this” — it is “was this recording generated at all,” and that is the question a detector can actually answer. Second, knowing you are dealing with Resemble AI audio changes the context: where the recording is likely to have come from, what the realistic risk is, and what you should do next. That context is what the rest of this page covers.

If you see a tool that confidently attributes a clip to a named vendor, it is fair to ask what evidence that claim rests on, and how it performs against models released after it was trained.

Context

Enterprise voice, and a vendor on both sides.

Where its audio shows up, and why it also sells the countermeasure.

Resemble AI provides voice cloning and speech synthesis aimed at organisations rather than individual creators, and it has also built products on the other side of the problem — detection of synthetic speech and watermarking of generated audio. Several companies in this space have moved the same way.

Where you actually encounter it

  • Branded and licensed voices. A single consistent voice across a company's IVR, products and content, often licensed from a real voice artist who was paid for it.
  • Localisation at scale. One voice identity carried across languages and markets.
  • Media and game production. Placeholder and production dialogue.
  • Accessibility. Including voice banking, where someone at risk of losing their speech preserves it — a use worth keeping in mind before treating all cloning as suspect.
Consent is the dividing line, not the technology. A cloned voice built with a paid, consenting artist under contract and a cloned voice built from someone's Instagram videos are the same technique and completely different acts. A detector cannot tell them apart, because the difference is not in the audio. That is worth remembering before reading a synthetic verdict as a finding of wrongdoing.
The bigger picture

Provenance beats detection, eventually.

The reason a synthesis vendor also builds watermarking is that blind detection has a structural ceiling, and everyone working on this knows it. A detector learns the traces left by the generators in its training data. New models leave different traces. Generation improves faster than detection corpora get rebuilt, so every deployed detector is permanently somewhat behind the newest systems. Care with your audio does not fix that; it is the honest limit of the whole category.

Provenance approaches invert the problem. Instead of examining audio for evidence of generation after the fact, they attach a signal at creation — an inaudible watermark, or signed metadata describing how a file was produced and edited. Industry efforts on content credentials and audio watermarking are both active, and several synthesis vendors now watermark their output by default.

Blind detectionProvenance / watermarking
Needs cooperation?No — works on any fileYes — only covers generators that participate
Unseen generatorsDegrades, sometimes badlySimply absent — no signal to find
Survives compression?PoorlyDepends on the scheme; a design goal, not a given
Adversary who caresCan post-process to evadeCan use a non-participating generator
Absence proves what?Nothing conclusiveNothing — no watermark is not evidence of authenticity

Both have gaps, and the last row is the one people misread most often. Neither an absent watermark nor a “human” detector verdict establishes that a recording is genuine, unedited or in context. They are different partial signals, and the sensible position is to use whatever you have and to keep provenance — where the file came from and who handled it — as the primary evidence.

Not forensic-grade. A result on suspected Resemble AI audio is a triage signal, not proof, and it must never be the sole basis for a disciplinary, employment, financial or legal decision. Our full position on what a result is worth is here.
Practical

How to check audio you think came from Resemble AI.

  1. Get the least-processed copy you can. Every forward, download and re-upload strips evidence. If the file reached you through several people, ask whoever had it first.
  2. Check the format. MP3, WAV, M4A, OGG and FLAC, up to 10 MB. The 10 MB limit is enforced on the server, so a larger file is rejected on every plan. WhatsApp voice notes arrive as Opus and need converting first — the WhatsApp guide covers it.
  3. Trim to clear, continuous single-speaker speech. Thirty seconds of one person talking cleanly beats five minutes of a noisy multi-speaker recording.
  4. Read the confidence, not just the verdict. A low-confidence “likely AI” and a high-confidence “likely AI” are different findings.
  5. Run a control clip. A recording you are confident is genuine, captured through a similar channel, tells you what “real” looks like in this situation. The comparison method is the highest-value move available to you.
Free, no signup: 3 checks a day. Signing in does not raise it on its own — voice runs on its own plans (Echo $9/mo for 100 checks, Amplify $29 for 1,000, Broadcast $99 for 10,000, or one-time Soundbites credit packs from $5), separate from our text subscriptions. Audio is processed and discarded — never stored, never used for training, never shared.
Other generators

Checking audio from a different tool?

The check is the same for all of them — the context around it is what changes.

ElevenLabsOpenAI TTSMurfPlayHT
FAQ

Resemble AI detection questions.

Can TextSight tell me a recording came from Resemble AI?
No. It returns an AI-likelihood score and a human-versus-synthetic verdict, not the identity of the generator. Source attribution from audio alone is substantially harder than binary detection and our detector does not attempt it.
Should I use a synthesis vendor's own detector instead?
Use it in addition if the audio might be theirs, since more independent signals is better. Understand the scope though: a vendor's detector is built around that vendor's own output, and it has nothing to say about audio from other models. A general detector has the opposite limitation. Neither is proof.
What is audio watermarking and does it help me?
It is an inaudible signal embedded at generation time so the audio can later be identified as synthetic. It helps when the generator participates and the watermark survives whatever compression the file went through. It does not help against a generator that does not watermark, and critically, the absence of a watermark is not evidence that a recording is genuine.
Is voice cloning legal?
Cloning your own voice, or someone's voice with their documented consent under a contract, is ordinary commercial practice. Cloning a real person without permission breaches the terms of every major provider and may breach personality, publicity or data protection rights depending on where you are. The technology is the same in both cases; consent is the difference, and no detector can see it.
A company cloned a voice for their IVR. Is that a problem?
Not inherently — licensed, consented brand voices are a normal commercial arrangement, and the artist is typically paid for it. Disclosure to callers is a separate question and one where obligations are tightening. Neither is something a detector can settle.
Can I use a detection result in a contract dispute?
Not as the basis for it. It is not forensic-grade and must never be the sole basis for a legal, commercial or employment decision. Where the stakes justify it, a qualified forensic audio examiner can do what an automated score cannot.
Related

More voice detection guides.

Checking an enterprise recording? Start here.

Three checks a day, free, no signup. Your audio is never stored.

Check an audio clip All voice guides
Human-vs-synthetic verdict · no engine attribution · not forensic-grade