HomeVoice Detection › PlayHT

Detect PlayHT AI voices.

PlayHT's centre of gravity is real-time conversational speech — the low-latency voice behind phone agents, support bots and interactive assistants that answer and talk back. That creates a problem no detector solves: by the time you want an answer you are already on the call, and detection needs a recorded file. This page is honest about that, and covers what you can actually do instead.

Check an audio clip free The live-call problem ↓
3 checks a day, no signup Audio never stored
Read this first

What we can and cannot tell you.

This is the part most pages about detecting PlayHT voices quietly skip.

Our detector does not name the generator. It returns an AI-likelihood score and a human-versus-synthetic verdict. It will not tell you “this was PlayHT” — that is source attribution, a substantially harder problem than telling synthetic from real, and our detector does not return it. We would rather show you nothing than a plausible-sounding guess.

So what is this page for? Two things. First, the useful question is almost never “which tool made this” — it is “was this recording generated at all,” and that is the question a detector can actually answer. Second, knowing you are dealing with PlayHT audio changes the context: where the recording is likely to have come from, what the realistic risk is, and what you should do next. That context is what the rest of this page covers.

If you see a tool that confidently attributes a clip to a named vendor, it is fair to ask what evidence that claim rests on, and how it performs against models released after it was trained.

Context

Built for agents that talk back.

The real-time use case is what makes this engine's situation different from the others.

PlayHT provides text-to-speech through an API with an emphasis on streaming and low latency — generating speech fast enough to hold a conversation rather than rendering a file to download later. That capability is what conversational voice agents are built on.

Where you actually encounter it

  • Inbound and outbound phone agents. Support lines, qualification calls, appointment reminders, collections.
  • In-product voice assistants. Speech inside an app or a website widget.
  • Interactive experiences. Anywhere a voice needs to respond to what you just said.
  • Standard narration too. The same API renders ordinary voiceover; real-time is the differentiator, not the only use.

The important consequence is that you are far more likely to meet this class of audio in a conversation than as a file someone sent you. And a conversation is precisely the situation a detector cannot help with.

The honest limit

You cannot check a live call.

There is no real-time detection running on your phone line. Not from us, and not from anyone else at consumer scale. Detection runs on a recorded file after the fact. If something is marketed as live protection, it is worth asking in detail what it actually does and where it runs.

So during the call itself, the things that work are behavioural rather than technical:

  1. Interrupt. Talk over the caller mid-sentence. Real conversation overlaps and recovers naturally; scripted and generated speech tends to restart a sentence, talk through you, or leave an unnatural gap.
  2. Go off-script. Ask something specific and unexpected that no flow could anticipate. An agent following a decision tree deflects, loops, or offers to transfer you.
  3. Ask directly. “Am I speaking to a person?” Legitimate deployments increasingly have to answer this honestly, and an evasive answer is itself informative.
  4. Never act on inbound urgency. Whatever the answer, if the call wants money, credentials or a code, hang up and call the organisation back on a number you look up yourself. The scam-call guide covers this properly.

Afterwards, if the call was recorded

If your phone records calls, or your organisation records inbound lines, you can check the recording. Set expectations first: telephone audio is the hardest case in this whole category because call compression is band-limited and aggressive, and it strips out much of the high-frequency detail detection depends on. Expect lower confidence than on a clean file, and treat an uncertain result as genuinely uncertain rather than as quiet confirmation.

This is exactly where the comparison method earns its keep. A recording of a known-human call through the same phone system gives you a baseline for what “real” looks like on that channel, and the gap between the two is worth far more than either score alone.

Not forensic-grade. A result on suspected PlayHT audio is a triage signal, not proof, and it must never be the sole basis for a disciplinary, employment, financial or legal decision. Our full position on what a result is worth is here.
Practical

How to check audio you think came from PlayHT.

  1. Get the least-processed copy you can. Every forward, download and re-upload strips evidence. If the file reached you through several people, ask whoever had it first.
  2. Check the format. MP3, WAV, M4A, OGG and FLAC, up to 10 MB. The 10 MB limit is enforced on the server, so a larger file is rejected on every plan. WhatsApp voice notes arrive as Opus and need converting first — the WhatsApp guide covers it.
  3. Trim to clear, continuous single-speaker speech. Thirty seconds of one person talking cleanly beats five minutes of a noisy multi-speaker recording.
  4. Read the confidence, not just the verdict. A low-confidence “likely AI” and a high-confidence “likely AI” are different findings.
  5. Run a control clip. A recording you are confident is genuine, captured through a similar channel, tells you what “real” looks like in this situation. The comparison method is the highest-value move available to you.
Free, no signup: 3 checks a day. Signing in does not raise it on its own — voice runs on its own plans (Echo $9/mo for 100 checks, Amplify $29 for 1,000, Broadcast $99 for 10,000, or one-time Soundbites credit packs from $5), separate from our text subscriptions. Audio is processed and discarded — never stored, never used for training, never shared.
Other generators

Checking audio from a different tool?

The check is the same for all of them — the context around it is what changes.

ElevenLabsOpenAI TTSMurfResemble AI
FAQ

PlayHT and voice agent questions.

Can I detect an AI voice during a live phone call?
No. Detection runs on a recorded audio file after the fact, not on a live stream, and that is true of every consumer tool we are aware of. During a call, interrupting, going off-script and refusing to act on inbound urgency are what protect you.
Can TextSight tell me a call used PlayHT?
No. It returns an AI-likelihood score and a human-versus-synthetic verdict, not the name of the engine. Source attribution from audio alone is a harder problem and our detector does not attempt it.
Why is call audio harder to analyse?
Telephone audio is band-limited and heavily compressed. The codecs are designed to discard whatever a listener will not consciously miss, which overlaps almost exactly with the fine spectral detail detection relies on. Expect meaningfully lower confidence than on an uncompressed recording.
Is a company allowed to use an AI voice to call me?
Rules on automated and artificial-voice calls vary by jurisdiction and several regimes require consent, disclosure or both — rules on artificial-voice calling in the US and transparency obligations in the EU are the commonly cited examples. If you believe you received an unlawful automated call, your national telecoms or consumer regulator is the place to raise it.
How do I record a call to check it later?
That depends on your device and, importantly, on your local law — consent requirements for recording a call vary and in some places recording without telling the other party is an offence. Check the rules that apply to you before relying on this. Many organisations already record inbound lines as a matter of course.
Does an uncertain result mean the caller was probably AI?
No, and this is the most common and most costly misreading. On compressed phone audio an uncertain result is usually a statement about the file rather than about the voice. Treat it as no information, not as weak support for what you already suspected.
Related

More voice detection guides.

Have a call recording? Check it afterwards.

Three checks a day, free, no signup. Your audio is never stored.

Check an audio clip All voice guides
Recorded files only · no live-call detection · phone audio is the hardest case