OpenAI's speech synthesis is different from the voice-cloning tools in one way that changes everything about the risk: it offers a set of preset voices rather than a clone of whoever you point it at. So you will not hear your own relative — you will hear a stranger who sounds completely natural, usually from inside some other product. That shifts what you should be worried about, and this page covers where.
This is the part most pages about detecting OpenAI voices quietly skip.
So what is this page for? Two things. First, the useful question is almost never “which tool made this” — it is “was this recording generated at all,” and that is the question a detector can actually answer. Second, knowing you are dealing with OpenAI audio changes the context: where the recording is likely to have come from, what the realistic risk is, and what you should do next. That context is what the rest of this page covers.
If you see a tool that confidently attributes a clip to a named vendor, it is fair to ask what evidence that claim rests on, and how it performs against models released after it was trained.
This is the one structural difference that matters, and almost nobody writing about it says so.
OpenAI's text-to-speech is an API product: developers send text and receive audio, using a set of provided voices. It is not built as a tool for cloning a specific person from a sample. That single design decision moves the risk from one place to another rather than removing it.
With preset-voice synthesis, the harm is rarely that someone was impersonated. It is that a person did not know they were talking to, or listening to, a machine — and made a decision they would not otherwise have made.
That is increasingly a regulated question rather than only an ethical one. Disclosure obligations for AI-generated audio and for automated callers exist in several jurisdictions and are tightening; the EU AI Act's transparency provisions are the most-cited example, and a number of US states regulate automated and artificial-voice calling directly. If you are deploying synthetic voice in a product, disclosure is now a compliance question for your legal team, not a matter of taste.
If you are on the receiving end and want to know whether an audio file was machine-generated, that is exactly what a detector answers — with the usual caveat that a recorded file is required. There is no way to run detection on a live call, from us or from anyone else at consumer scale.
The check is the same for all of them — the context around it is what changes.
Why we publish no single accuracy percentage, what false positives look like, and how to weigh a result.
See our position →Run a known-genuine recording alongside the suspect one and read the gap. The strongest method available.
Learn the method →The technical guide: how detection works, and why it generalises poorly to generators it has not seen.
Read the technical guide →Three checks a day, free, no signup. Your audio is never stored.