AI writing tools invent citations. Not rarely, and not obviously: the fakes carry real author names, plausible journals and well-formed identifiers. This is a free book about why that happens, the three shapes a fake reference takes, and the database method that settles any citation in about a minute. Free to read here, free to share, no email required.
A citation has one job. It has to point at a real source that says what you say it says, in a form your reader can go and find. AI writing tools break that job quietly. They produce references with the exact shape of real ones, and a share of those references point at nothing at all.
This book is about that problem and the fix, which is duller and more reliable than most people expect. There is no detector that settles a citation. There is a method, it takes about a minute per reference, and it works every time.
It is written for anyone who has to stand behind what they cite: students, researchers, editors, and lawyers, who have the sharpest version of this problem and the most public consequences when it goes wrong. If you remember one sentence from the whole book, make it this one. Never cite a reference an AI tool handed you without confirming it yourself.
A large language model is not a search engine and not a library. It predicts the next most likely piece of text based on patterns it learned during training. When you ask for a citation, it produces something with the statistical shape of a real reference, the right author cadence, a plausible journal name, a believable year, without ever checking whether that exact paper was published.
So the honest summary is this. ChatGPT can quote real papers it saw often enough during training. It can also stitch together references that never existed. And it usually cannot tell you which is which, because to the model both came from the same process. That is why "ChatGPT gave it to me" is never a citation. It is a draft you still have to verify.
The exact rate moves with the model, the prompt, and how obscure the topic is, so treat any single percentage with caution. What the research consistently shows is that the problem is common, not rare.
Studies of AI-generated bibliographies have found that roughly 15 to 30 percent of citations are fully fabricated, meaning the paper does not exist. On top of that, another 20 to 30 percent or so contain errors even when a real source sits underneath: a wrong author, a mismatched journal, an incorrect year, or a DOI that resolves to a different article. A widely cited Nature Scientific Reports analysis of ChatGPT bibliographic references documented this pattern of invented and inaccurate citations.
Two takeaways follow. First, fabrication is not an edge case you can ignore. If you pull ten references from a chat model, expect several to be wrong in some way. Second, "wrong" is broader than "invented." A reference can be half true, which is harder to catch because parts of it check out. We get into the precise breakdown of how reliable these references really are in Are AI citations reliable?
It helps to understand the mechanism, because once you see why fabrication happens you stop being surprised by it and start checking by default.
A model like ChatGPT was trained to continue text in a plausible way. It has no live connection to CrossRef, PubMed, or any journal index when it answers. It works from a compressed memory of patterns in its training data, not a lookup table of real papers. So when you ask for "three peer-reviewed studies on X," it generates text that reads like three peer-reviewed studies. Author names that go together, a journal that publishes in that field, a year that fits, a DOI in the right format. The output is optimized to be convincing, not true.
This is why the fakes are so polished. The model is good at the surface form of a citation and has no check on the underlying fact. The same machinery that lets it write a fluent paragraph lets it write a fluent reference to a paper nobody ever wrote. We unpack this in full, including why newer "browsing" features help but do not fully solve it, in Why does ChatGPT make up citations?
It is worth saying plainly: this is not the model lying. It is doing what it was built to do, predict plausible text. The fix is not to expect the model to stop. It is to verify what it produces.
Not all fake citations fail the same way, and knowing the categories makes them easier to catch. There are three common shapes.
The cleanest case to detect, once you look. The paper does not exist at all. No database has it, no search turns it up, the DOI resolves to nothing. The author names may be real researchers in the field, which is what makes the reference feel safe, but the title and publication were generated whole.
The trickiest of the three. A chimera reference is assembled from real parts that never belonged together: a genuine author, a real journal, an actual paper title, combined in a way that does not match any single published work. The author wrote real papers, just not this one. The journal is real, but it never ran that article. Because every piece checks out, a quick glance fails to catch it. Only following the reference all the way to the actual source exposes it.
A real paper underneath, but with the details mangled. The year is off by two. The page numbers are wrong. The journal name is close but not exact. The DOI points to a different article. These are dangerous because the source is real, so a casual check ("does this paper exist?") passes, while the claim you attribute to it may not be supported by the actual text.
Recognizing which type you are dealing with shapes how you verify. We walk through detection tactics for each in Types of fake AI citations explained.
Before you verify against a database, a few tells can flag a reference as suspicious in seconds. None is proof on its own, but together they tell you where to look harder.
This is the part that actually settles it. Verifying a citation is not hard, it just takes a minute per reference and a few free tools. There is no shortcut that skips the databases, and anyone who claims a tool can confirm a citation without checking it against a source index is overselling.
Here is the manual method, in order:
Run those four steps and you will catch the overwhelming majority of fabricated, chimera, and distorted references. We expand this into a full cross-model workflow, including references from Claude, Gemini, and Perplexity, in How to verify AI-generated references, and cover the fast single-reference check in How to check if a ChatGPT citation is real. If you have already found a fake, ChatGPT made up a source: what to do walks through the cleanup.
The stakes scale with the setting, and three groups feel them most.
Students. A single fabricated reference in an essay can read as a fabricated source to an instructor, and fabrication is treated seriously by most academic integrity codes regardless of intent. "The AI gave it to me" is not usually accepted as a defense. We cover where the line sits, and how reviewers catch this, in Are fake AI citations academic misconduct?
Researchers and editors. A fake or distorted citation that slips into a manuscript can survive peer review, mislead readers, and damage credibility once caught. Editors increasingly spot-check references, and a fabricated source can get a paper rejected or retracted.
Lawyers. This is where it has gone furthest, with real professional sanctions, which brings us to the case everyone in law now knows.
The clearest cautionary tale is Mata v. Avianca. In June 2023, two New York attorneys submitted a federal court filing that cited multiple cases which did not exist. The fake case law had been produced by ChatGPT, and the lawyers had not verified it. Judge P. Kevin Castel of the Southern District of New York sanctioned the attorneys and their firm, ordering a $5,000 penalty and requiring them to notify the judges whose names had been falsely attached to the invented opinions.
That case was not the last. Courts in several jurisdictions have since sanctioned attorneys for filings built on AI-fabricated citations, and judges have made clear that the duty to verify rests with the person who signs the document, not the tool. The lesson generalizes beyond law: an AI citation you have not checked is a liability you sign your name to. We go deeper on the cases and the professional duty in Fake AI citations in legal briefs.
ChatGPT can and does generate fake citations, in three recognizable forms, because it predicts plausible text rather than retrieving real records. That is not a bug you can prompt your way around. It is how the tool works. The good news is that verification is straightforward: resolve the DOI, search the exact title, cross-check a real database, and confirm the source supports your claim. Do that for every reference and you keep the speed of AI drafting without inheriting its fabrications.
Start with the mechanism, because once you understand why fabrication happens you stop being surprised by it and start checking by default.
The most useful mental model is this: ChatGPT is a very advanced autocomplete. It was trained on an enormous amount of text, and from that it learned the statistical patterns of language, which words and phrases tend to follow which others. When you give it a prompt, it produces a response one piece at a time, each time picking a word that is likely to come next given everything before it.
Notice what is missing from that description. There is no lookup step. The model is not consulting a stored copy of the source it names. It is not pinging Google Scholar or CrossRef in the background. It is doing the same thing for a citation that it does for a sentence about the weather: producing the most plausible continuation it can.
This is why people in the field describe a base chat model as having no built-in fact-checker. It is optimized to sound right, not to be right. Most of the time, sounding right and being right line up, because the patterns it learned came from accurate text. With citations, they come apart more often, and that gap is where fake sources are born.
You have probably seen the word hallucination used for these errors. It is the standard term for when an AI model states something that is plausible and confident but false. A hallucinated citation is just a hallucination that happens to be shaped like a reference.
The word is a little misleading, because it suggests the model malfunctioned. It did not. Generating a smooth, confident, wrong answer is the system working exactly as designed. The model's whole job is to produce fluent, likely-sounding text, and a fabricated citation is extremely likely-sounding text. It has the right format, the right rhythm, the right academic register. Nothing in the model's training tells it to stop and ask whether the specific paper it just named is real.
That is also why hallucinated citations feel so trustworthy. They are not sloppy or obviously broken. They are fluent by construction, which is the very thing that makes them dangerous to anyone who pastes them in without checking. We break the failures into categories on our page covering the different types of fake AI citations, from fully invented references to ones stitched together from real parts.
Plenty of AI output is just prose, and prose hallucinations are usually easier to catch because a wrong claim often sounds a little off. Citations are a special case for a few reasons, and they stack on top of each other.
A reference has a rigid, predictable shape. Author, year, title, journal, volume, pages, DOI. The model has seen millions of them, so it is exceptionally good at producing the form. But the form is the easy part. The form does not require the underlying paper to exist. The model can fill every slot with something believable and never touch reality.
Author names, journal titles, and publication years are common, reusable tokens. The model knows that papers about memory often appear in journals with words like Cognition or Neuroscience in the title. It knows the years that look current. It knows what a DOI prefix looks like. So it can assemble a reference out of plausible parts the same way it assembles a sentence out of plausible words, and the parts can be individually real while the combination is fiction.
When you ask a person for a source and they do not know one, they can say so. A base language model has no reliable internal signal that says, "I do not actually have this reference." Producing a confident citation and producing the words "I do not know" are both just sequences of tokens to it, and a confident, complete-looking answer usually matches the training data better. So it tends to fill the gap rather than admit it.
Suppose you ask for a source on how sleep affects memory and you get something like this:
Hartman, L., & Boyd, R. (2019). Sleep consolidation and long-term memory retention in adults. Journal of Cognitive Neuroscience, 31(4), 612-627. https://doi.org/10.1162/jocn_a_01388
Look at how complete it is. The author names are ordinary. The year is reasonable. Journal of Cognitive Neuroscience is a real journal, and the topic genuinely fits it. The volume and page numbers are formatted correctly. The DOI even uses the real prefix that journal's publisher uses. Every individual element passes a glance.
And yet a reference that looks exactly like this can be entirely manufactured. The authors may never have written it together, the article title may not exist, and the DOI may resolve to a different paper or to nothing at all. The model did not copy a real citation. It produced the most plausible-looking reference for your prompt, slot by slot. That is the whole problem in one example: plausibility was the goal, and existence was never checked.
This is also why you cannot tell a fake citation from a real one just by reading it. Fabricated references are built to pass the eye test. The only way to know is to verify the actual source, which is a separate skill we treat on its own.
A fair question: surely GPT-4, and the models after it, fixed this? They are better, but they have not eliminated it, and the reason follows directly from the mechanism above.
Scaling a model up makes it a better predictor of plausible text. It does not change what the model is fundamentally doing. A larger model produces references that look even more convincing, with cleaner formatting and more topically perfect journal choices. In one sense that is worse, because the fakes get harder to spot by eye even as they get somewhat rarer. Research on AI citation fabrication has found meaningful drops in invented references between model generations, but not a drop to zero, and the share that is wrong or distorted in some way stays stubbornly present. Treat any single percentage you see as a rough, source-dependent estimate rather than a fixed law, since the numbers vary a lot by study and by how strictly "fake" is defined.
There is one real fix worth naming, because it changes the picture. Some AI tools now connect the model to an actual search step. When a system browses the web or pulls from a real citation database and then writes its answer from those retrieved results, the citations can point to documents that genuinely exist. That is a different setup from a plain chat model answering from memory, and it is more reliable for sources. But two cautions apply. First, a tool that says it can browse does not always browse for every answer, and it can still summarize a real source inaccurately. Second, the default behavior of a standalone chat session, with no live retrieval, is exactly the next-token prediction we described, which is why the plain-model case is the one to plan around. Whether any of this makes AI citations trustworthy enough to rely on is its own question, and we take it up in are AI citations reliable.
The practical takeaway is narrow and clear. If a citation came out of a language model, you cannot assume the source exists, even when it looks flawless, even from a top-tier model, even when the journal and the topic match perfectly. The fluency is not evidence. It is the symptom.
That does not make AI useless for research. It is a strong drafting and brainstorming aid. It just means a citation from it is a lead to check, not a finished reference to paste. The reliable move is to confirm every source against the real record before it goes into your work: search the title, resolve the DOI at the publisher, and look the paper up in a database like Google Scholar, CrossRef, or PubMed. If the paper does not turn up where it should, that is your answer.
Knowing why it happens does not tell you how often. That is the next question, and it is the one that decides how much checking your own work actually needs.
A citation has one job. It has to point to a real source that says what you claim it says, in a form a reader can find. That is a higher bar than it sounds, because a citation can fail in several quiet ways:
A reference can be a fabrication, a near-miss, or a real source twisted to fit. "Reliable" means it survives all three tests. Most AI citations are not tested against any of them before they reach a document, which is the whole problem.
The honest version of the numbers is a range, not a single figure, and it depends on the model, the field, and how the question was asked. But across the studies that have looked at this, a consistent picture emerges.
Studies have found that roughly 15 to 30 percent of citations generated by general-purpose chatbots are fully fabricated, meaning the source does not exist at all. On top of that, another 20 to 30 percent of the citations that do point to real sources contain errors in the details, such as the wrong authors, year, page numbers, or DOI. A peer-reviewed paper in Nature's Scientific Reports examined ChatGPT's bibliographic citations and found a substantial fabrication rate, with the references it produced frequently failing verification against the real literature.
Two caveats keep this honest. First, those figures come from specific tests at specific times, often on earlier model versions, so the exact percentages move as models change. Second, newer models and retrieval-grounded tools tend to do better than the worst early results. The trend is improving. But "improving" is not the same as "reliable," and even a 10 percent fabrication rate is catastrophic if the one fabricated citation is the one a reviewer checks.
So the accurate summary is this: a noticeable fraction of AI citations are invented, a further noticeable fraction are wrong in their particulars, and you cannot tell which is which by reading them. That last part is what makes the reliability question impossible to answer with "trust them."
This is the core of it. A plain chatbot like the default ChatGPT, Claude, or Gemini, asked to produce citations from its own knowledge, is the worst case for reliability, and it is worth understanding why so the rule sticks.
A language model does not look anything up when it writes a citation from memory. It predicts the most likely next words. A citation has a very regular shape: author, year, title, journal, volume, pages. The model has seen millions of them, so it can generate something that has the exact texture of a real reference, with a believable author for the field and a journal that publishes that kind of work. None of that requires the source to exist. The output is shaped to look right, not to be right.
That is the mechanism behind the fabrication numbers, and we explain it in more depth in why does ChatGPT make up citations. The practical takeaway for the trust decision is simple. When a citation comes from a model writing from its training rather than from a real document in front of it, treat the reference as a hypothesis, not a fact. It might be real. You have no way of knowing until you check, and the base rates above mean a real chunk of them will not survive the check.
Not all AI tools work the same way, and the difference matters for reliability.
Search-grounded and retrieval-augmented tools, such as Perplexity, the search or browsing modes in ChatGPT and Gemini, and research assistants that pull from real databases, do something the plain chatbot does not. They retrieve actual documents first and then write with those documents in view. Because there is a real source behind the link, the citation is far more likely to point to something that exists. For "does this source exist," these tools are meaningfully more reliable.
But better is not the same as safe, and the failure modes shift rather than disappear:
So retrieval-grounded tools move the needle from "frequently invents sources" to "usually points at a real source it may have misread." That is a real improvement and worth preferring. It is not a green light. You still have to open the source and confirm it says what you are about to claim it says.
Putting it together, here is the trust decision in plain terms.
You can lean on an AI citation a little more when all of these are true:
In other words, the only AI citation you can lean on is one you have already verified, at which point you are trusting the source, not the AI. The AI just found a candidate for you.
You must not lean on an AI citation when any of these are true:
The line is not "chatbots bad, search tools good." The line is verified versus unverified. A search-grounded link you never opened is still unverified. A plain-chatbot citation you fully checked against the real database is fine, because you did the work the AI did not.
Because the whole reliability question collapses into "did you verify it," the verification step has to be quick enough that you will actually do it. The manual method takes a minute or two per citation and uses tools that are free and authoritative.
For the full cross-model workflow, including how to handle references from Claude, Gemini, and Perplexity, see how to verify AI-generated references. The point of the manual method is that it does not rely on any tool's promise. You are checking the citation against the actual scholarly record, which is the only thing that settles the question.
AI citations are not reliable on their own. A real share of them are invented, a further share are wrong in the details, and you cannot tell the good ones from the bad ones by reading them. Plain chatbots writing from memory are the least trustworthy source of references there is. Retrieval and search-grounded tools are better because there is a real document behind the link, but they still misread sources and still mix in generated guesses, so "better" never reaches "no need to check."
The rule that survives all of this is short. Never submit an AI citation you have not verified yourself. Treat every reference an AI gives you as a lead to be checked, not a fact to be trusted, and the reliability question takes care of itself. The opening chapter covers the bigger picture of why models invent references at all.
Not every fake fails in the same way, and the differences matter. One kind falls apart the moment you look. Another survives every quick check you are likely to run.
When people first hear that AI fabricates references, they picture a citation conjured out of thin air: a paper that does not exist by authors who do not exist. That happens, and it is the easiest type to disprove. The other two look far more legitimate. They borrow real names, real journals, and real digits, so a quick glance or casual search can wave them through.
Researchers who catalogue AI citation errors tend to sort them into roughly these buckets: fully fabricated references, references where real and false elements are mixed, and references that point at a real work but get the metadata wrong. The proportions shift by model and topic, but published testing has put fully fabricated citations somewhere in the range of 15 to 30 percent of AI-generated references, with a further fifth to a third carrying factual errors of the kind described below. Treat those as attributed ranges from the literature, not fixed constants. Errors and outright inventions together make up a large share, and they do not all hide in the same place.
This is the textbook fabrication. The paper does not exist. The authors, as a writing team, do not exist. The journal issue is empty, or the journal itself is made up. The DOI, if any, resolves to nothing. All of it was generated.
Hartwell, J. R., & Okonkwo, A. (2021). Cognitive load and retrieval fatigue in adaptive learning environments. Journal of Applied Cognitive Research, 14(3), 221-244. https://doi.org/10.1080/14792779.2021.1934567
It reads perfectly. The title is plausible, the format is clean, the DOI looks the part. But search the title in quotation marks and nothing comes back. The journal name matches no real publication. The DOI prefix belongs to a real publisher, yet that exact suffix resolves to an error page. None of this reference exists.
A language model does not look anything up. It predicts the next most plausible string of text from patterns it has seen. When you ask for a citation, it produces something with the shape of one: a name in the right place, a year, a title that fits the topic, a DOI in the right format. None of that involves checking a database, because the model has none. So when the training data holds no real source for your request, it generates one that looks right. For the deeper reason a model invents a reference rather than admitting it has none, The chapter on why a model invents sources covers that mechanism in more depth.
The fully invented citation usually has no footprint anywhere. The clearest tells:
A simple search defeats this type. Paste the exact title into Google Scholar or a library database inside quotation marks. If a real paper has that title, it appears. If nothing appears, that is your answer. Then resolve the DOI at doi.org. A real DOI lands on the actual article page; a fabricated one returns a "DOI Not Found" error. The DOI chapter walks through that resolution step, because a dead DOI is the single fastest tell for this whole category.
The chimera is the dangerous one. Every individual element is real. The author is a real researcher who really does publish in that journal, and the year is one when that author was active. But the specific paper, that title by that author in that journal in that year, does not exist. The model assembled true parts into a false whole.
Tversky, A., & Kahneman, D. (1979). Heuristic anchoring under conditions of statistical uncertainty. Cognitive Psychology, 11(2), 168-192.
Tversky and Kahneman are real, foundational researchers who genuinely co-authored landmark work. Cognitive Psychology is a real journal that has published cognition research for decades. 1979 sits inside their active collaboration, and anchoring is a real concept they studied. Every part checks out individually. But that exact paper, with that exact title, was never published. The reference is a plausible Frankenstein of authentic limbs.
The model has seen these authors, this journal, and this topic appear together many times in training, so it has learned the association strongly. When you ask for a source on anchoring, it stitches together the names, the venue, and a title that sounds like something they would have written. Because all the components co-occur naturally, the fabricated combination feels more convincing than a fully invented one. The model is, in effect, pattern-matching its way to a paper that should exist by the logic of the field but does not.
This is where casual verification fails. Search the author and you find them. Search the journal and you find it. Each isolated check passes, which is why people stop checking. The tells are subtler:
You have to verify the combination, not the parts. Searching the author alone or the journal alone will mislead you, because both are genuine. Instead, search the exact title in quotation marks and confirm that specific paper exists. Then cross-check against the author's authoritative record: their ORCID page, Google Scholar profile, or institutional listing. If the paper is real, it appears in their own bibliography. If it does not, the combination is fabricated no matter how real the ingredients are. Finally, resolve the DOI if one is given, and confirm the landing page matches the title, authors, and year. A walkthrough for this cross-check lives in how to check if a ChatGPT citation is real. The chimera is the reason "I checked the author and they're real" is not enough.
The distorted citation is a real paper with a detail wrong. The source genuinely exists, and you can find it. But something has been altered: the year is off by one, a co-author is misspelled or dropped, two DOI digits are transposed, or the page numbers do not match. The bones are authentic; the metadata is corrupted.
Suppose the real paper is:
Bandura, A. (1977). Self-efficacy: Toward a unifying theory of behavioral change. Psychological Review, 84(2), 191-215.
A distorted version might read:
Bandura, A. (1978). Self-efficacy: Towards a unifying theory of behavioural change. Psychological Review, 84(3), 191-215.
The paper is real and famous. But the year is wrong by one, the issue number has shifted from 2 to 3, and small wording changes have crept into the title. A reader who recognizes the work might nod it through, because the source clearly exists. Yet the citation as written points at something that does not match the record.
Distortion usually comes from the model reproducing a genuine source from memory and getting the fine details slightly wrong, the way a person quoting a half-remembered reference might shift a date or a page number. It can also come from blending two nearby real papers, or from a DOI where the model predicted plausible digits and landed one or two off. Because the underlying work is real, this type is the least likely to be flagged as "fake" and the most likely to slip into a finished bibliography, which makes it quietly corrosive.
The source exists, so a title search succeeds and can lull you into stopping. The tells are in the mismatch:
Resolve the DOI, then compare every field against the page it lands on. This is the step people skip: they confirm the DOI resolves and assume the job is done. Resolution alone is not enough for distorted citations, because a transposed digit can resolve to a real but different paper, and an accurate-looking DOI can sit beside a wrong year in the citation. Read the destination page and check the title, every author, the year, the volume, the issue, and the page numbers against what the reference states. If any field disagrees, the citation is distorted and needs correcting. The DOI-comparison technique is laid out step by step in how to check if a DOI is real or fake.
The quickest way to internalize the taxonomy is to notice that each type survives a different level of checking:
That progression is why a single habit, "I searched the author and they're real," catches the first type, misses the second, and never even tests the third. Thorough verification has to go all the way to comparing the specific paper, field by field, against an authoritative database. If you are weighing how much to trust AI-supplied references at all, are AI citations reliable? looks at what the testing actually says about their trustworthiness.
A taxonomy of fake citations is a verification problem, and verification is something you do by hand against the real databases: doi.org, CrossRef, Google Scholar, PubMed, OpenAlex, and ORCID. No tool replaces that comparison. Be wary of any product that claims to confirm a citation is real with one click, because confirming a citation means matching it against the source record, the manual work above.
Now to the checking itself. Start with the fastest test there is, because on a good day it settles a reference in about ten seconds.
DOI stands for Digital Object Identifier. It is a permanent address for a piece of published work, most often a journal article, and it is meant to keep pointing at that work even if the publisher moves the file or redesigns its website. Think of it as a serial number for a paper rather than a link that can rot.
Every DOI follows the same shape. It starts with 10., then a prefix that identifies the registrant (the publisher or organization that registered it), a slash, and then a suffix the registrant chose. So you see things like 10.1038/s41586-020-2649-2 or 10.1037/0003-066X.59.1.29. The 10. never changes. The four-or-five-digit prefix after it (1038, 1037, 1016, and so on) maps to a specific publisher: 1038 is Nature, 1016 is Elsevier, 1037 is the American Psychological Association.
This structure is exactly why fake DOIs are so common and so catchable. The format is simple enough that a language model can produce something that looks correct, and rigid enough that a real registry either recognizes it or it does not.
A real DOI resolves. That is the whole test.
To resolve a DOI, you send it to the official resolver. There are two easy ways to do it:
https://doi.org/ in front of the DOI and open it in your browser. So 10.1038/s41586-020-2649-2 becomes https://doi.org/10.1038/s41586-020-2649-2. If the DOI is real, your browser lands on the publisher's page for that exact article.You can do this from your phone in a waiting room. There is no account, no paywall on the resolution step itself, and no tool to install. The DOI either takes you to a paper or it returns an error.
Now read what you landed on. This is the part people skip, and it is where the interesting failures live.
The resolver opens the publisher's page, and the title, authors, and journal on that page match the citation you were checking. This is the clean pass. The identifier is real and it points at the work the citation claims. You are not done verifying the whole citation, because the surrounding details can still be wrong, but the DOI itself is genuine.
You get an error like "DOI Not Found" or a page from doi.org saying the handle does not exist. This is the strongest single signal of a fabricated citation. A DOI that was never registered cannot resolve, and AI tools regularly produce DOIs that were never registered, because the model is generating a plausible-looking string rather than retrieving a record that exists. A dead DOI on a citation you cannot otherwise locate is, in practice, the closest thing to a confession that the reference was invented.
One caution before you condemn it: a very recently published article can occasionally have a DOI that has not finished propagating, and a typo in transcription can break an otherwise real DOI. So when a DOI fails to resolve, copy it carefully one more time and try the title search described below. If both come up empty, you are almost certainly looking at a fabrication.
This is the sneaky one. The DOI resolves cleanly, but the paper it opens is not the paper in your citation. The title is different, the authors are different, or the year is off by a decade. This is a sign of a chimera citation, a reference stitched together from parts of real records that do not actually belong together. The DOI is real because it was copied from some genuine article, but it was attached to a title and author list that the model assembled separately. We break down that pattern, along with fully invented and distorted citations, in types of fake AI citations.
A chimera is dangerous precisely because the DOI passes the resolve test. If you only checked that the link works and never compared the destination to the citation, you would wave it through. So the rule is: resolving is necessary, but you have to look at what it resolved to.
doi.org tells you whether a DOI resolves. CrossRef tells you what the registered metadata for that DOI actually says, which is what you need to catch a chimera.
CrossRef is the registration agency for most scholarly DOIs, and its public search at search.crossref.org lets you look up a DOI or a title for free. Two ways to use it:
CrossRef is not the only registry. Some DOIs, especially for datasets, theses, and preprints, are registered through DataCite rather than CrossRef, so a scholarly-but-not-journal DOI that is missing from CrossRef is not automatically fake. Resolving it at doi.org is still the deciding test for whether it exists at all.
For a fuller walkthrough that puts the DOI check inside a complete citation review, including title and author verification, see how to check if a ChatGPT citation is real.
You can verify a citation many ways. You can search the title, hunt for the authors, check whether the journal exists, confirm the volume and page numbers. All of that works, and all of it takes longer than resolving a DOI.
The DOI check is fast because it is binary and it is centralized. There is one official resolver, the answer comes back in seconds, and a registered DOI cannot be faked into existence. Contrast that with a title, which a model can copy verbatim from a real paper while changing everything around it. The DOI is the hardest part of a citation for an AI to fake convincingly and the easiest part for you to test, which is why it belongs at the front of any verification routine. If the DOI dies, you often do not need to check anything else.
This is also why the absence of a DOI is its own small signal. Most legitimate journal articles from the last fifteen years have one. A citation to a recent paper that conspicuously lacks a DOI is worth a closer look, though plenty of older work, books, and conference papers genuinely have none.
AI-fabricated DOIs are not random gibberish, and that is what makes them slip past a quick glance. A few patterns to watch for:
10.1038 or 10.1016, because it has seen thousands of them, then attaches an invented suffix. The string looks completely legitimate. It still will not resolve, because the full identifier was never registered. Always resolve the whole DOI, not just eyeball the prefix.doi.org path with the wrong format. If it does not start with 10. and follow the prefix-slash-suffix shape, it is not a DOI, and that alone tells you the citation was not assembled from a real record.None of these survive the two-step routine: resolve at doi.org, then compare the CrossRef metadata to the citation.
It is worth being clear about the difference between verifying a citation and detecting AI-written text, because they answer different questions. Resolving a DOI tells you whether a specific reference is real. It does not tell you whether the surrounding prose was drafted by a model. Those are separate checks for separate problems.
The chapter on why a model invents sources covers why this happens in the first place.
A DOI settles some references and not others. Plenty of citations arrive without one, and a chimera can carry a DOI that resolves perfectly to the wrong paper. So here is the full check, in order, for a single reference you need to be sure about.
Often enough that you should check every time. Published studies on ChatGPT's bibliographic output have found that roughly 15 to 30 percent of generated citations are fully fabricated, with another portion, somewhere in the 20 to 30 percent range, containing real-looking but incorrect details such as a wrong page range, a mismatched author, or a DOI that points somewhere else. Those numbers move around by model version and by field, so treat them as a ballpark rather than a fixed law. The practical takeaway does not change: a meaningful share of any ChatGPT reference list will not survive a check, so checking is not optional.
The good news is that verification is mechanical. You do not need judgment or expertise to do it. You need three free databases and about two minutes per source.
Start with the title, because the title is the fastest pass or fail.
Copy the exact paper title from the citation and paste it into Google Scholar inside quotation marks, like "the exact title here". The quotes force an exact-phrase search, so you are asking a direct question: does a paper with this precise title exist?
Three things can happen.
If you prefer a cleaner search experience, you can run the same exact-title query in Semantic Scholar or OpenAlex, both of which are free and index tens of millions of papers. They are good cross-checks when Google Scholar is ambiguous, and OpenAlex in particular is generous with metadata you can compare against.
Most ChatGPT citations include a DOI, the string that usually starts with 10. and acts as a permanent address for a paper. A DOI is the easiest single field to verify, because a real DOI always resolves to a real page.
Take the DOI from the citation and paste it onto the end of https://doi.org/, so it reads https://doi.org/10.xxxx/xxxxx, then open it. The DOI resolver will redirect you to the publisher's page for that exact paper.
One caution: a missing DOI is not proof of fabrication on its own. Older papers, books, and some conference proceedings legitimately have no DOI. If there is no DOI, you simply lean harder on the title search and the metadata cross-check instead.
By now you may already know the answer, but for anything you intend to actually cite, finish the job by confirming that the pieces belong together. This is the step that catches chimera references, the ones where every individual part is real but the combination is invented.
Open CrossRef, which is the registration agency behind most academic DOIs and lets you search by title, author, or DOI for free. If your source is medical or life-sciences, use PubMed instead, which indexes biomedical literature and is authoritative for that field. Search the title or the DOI, pull up the official record, and compare it field by field against the citation ChatGPT gave you:
If every field agrees with the official record, you are done. The citation is real and accurate.
Across those three steps, every citation falls into one of three buckets. Holding this simple model in your head keeps you from over-thinking a clear case.
The title returns an exact hit, the DOI resolves to that same paper, and the author, journal, and year all match the official database record. Everything points to one real document. You can cite it with confidence. In practice the strongest "Verified" is when the DOI resolution and the CrossRef or PubMed record independently agree with the title search, because that means two separate systems confirmed the same paper.
Some pieces are real and some are not. The author exists but the title is nowhere to be found. The DOI resolves but to a different paper. The paper is real but it ran in a different journal or year than the citation claims. Partial matches are the dangerous middle, because they look credible at a glance and often survive a lazy check. Treat a partial match as a fail. Do not cite it as written. Either find the correct details from the official record and fix the citation, or replace the source entirely.
The exact title returns nothing across Google Scholar, Semantic Scholar, and OpenAlex, the DOI is dead or absent, and no database has a record. This is a fabricated reference. There is no paper to fix. Delete it and find a real source that supports the same point.
It helps to know why a citation fails, because the reason tells you what to do next.
For a deeper breakdown of these categories and how to spot each one, the cross-model companion guide on how to verify AI-generated references walks through the same checks applied to Claude, Gemini, and Perplexity output, which behave a little differently from ChatGPT.
Once you have done a few, the whole thing compresses into a habit:
Do this for every reference ChatGPT hands you, not just the ones that look suspicious. Fabricated citations are designed, by the nature of how the model works, to look exactly as plausible as the real ones. The formatting is perfect on both. The only difference is whether the paper exists, and the only way to know is to look.
To be clear about what we do and do not do: TextSight does not have a one-click button that resolves every DOI in a document for you. Verifying a specific citation is the manual, database-driven process above, and we would rather teach it honestly than overpromise.
What we do offer is a hallucination detector that flags passages of AI output that read as fabricated or unsupported, which is a useful first pass for spotting where in a long ChatGPT response the invented material is likely to be, so you know which claims and citations to scrutinize first. Use it to narrow your attention, then run the three-step verification above on each citation it draws your eye to. The detector points; the database method confirms.
One citation is a minute. Thirty citations is an afternoon, and that is where good intentions quietly fail. This chapter is about doing it at scale without cutting the corner that matters.
A chat model does not keep a library. When it produces a citation, it is assembling the most statistically likely arrangement of a title, an author, and a journal, not retrieving a record from a database. That is why the output can be a "chimera": a real author bolted to a real journal under a title that does not exist, with a DOI that resolves to something else entirely or to nothing.
If you want the deeper explanation of why this happens and the patterns that give it away, the pillar guide can ChatGPT generate fake citations covers the mechanism in full. The takeaway is just this: a citation is a claim that a specific document exists and says a specific thing. Verification is the act of confirming both halves.
Not all assistants behave the same way, and this is worth understanding before you start.
Here is the catch. Retrieval-grounded models like Perplexity are more reliable about the link existing, but that does not mean the source supports the claim, that the page is authoritative, or that a cited study reports what the answer says it does. A real link to a real blog post is still the wrong citation for a clinical claim. So a search-grounded model lowers the fabrication rate, it does not remove the need to check. Verify retrieval-model references too, just with more attention to relevance and authority than to existence.
Verification is only as good as the index you check against. These are the free, authoritative ones, with the field each is best suited to. Bookmark them; this is your toolkit.
CrossRef is the registration agency behind most scholarly DOIs and holds roughly 160 million records across nearly every discipline. Its free search at search.crossref.org is the single most useful starting point because almost any legitimate journal article, conference paper, or book chapter from a participating publisher will be in it. If a reference claims a DOI, CrossRef is where you confirm the DOI belongs to that exact title and author. If a title that should be indexed returns nothing in CrossRef, that is a strong fabrication signal.
Every real DOI resolves. Paste it after doi.org/ in your browser, or enter it at the DOI resolver, and a genuine identifier sends you to the publisher's landing page for that work. A DOI that returns "DOI Not Found" is the quickest tell of an invented citation, and it takes about five seconds. We break this down step by step in how to check if a DOI is real. The one trap: a DOI can resolve to a real paper that is not the paper the citation describes, so always read the landing page, do not just confirm that it loaded.
For anything clinical, biomedical, nursing, pharmacology, or life sciences, PubMed is the authority. It indexes more than 37 million citations from MEDLINE and life-science journals, and its PMID is a second, biomedical-specific identifier you can confirm alongside the DOI. If a health-related reference is real, it is almost certainly in PubMed. If it is not, be skeptical.
OpenAlex is a free, open catalog of hundreds of millions of scholarly works, authors, and venues. It is excellent for cross-checking metadata (does this author exist, did they publish in this venue, is the year plausible) and for fields where you want wide coverage without a paywall. Because it is open and queryable, it is the database of choice when you want to confirm that the pieces of a citation actually belong together.
Semantic Scholar indexes roughly 200 million papers and is particularly strong in computer science, engineering, and AI, the fields where chat models are, ironically, most likely to be asked for references. It is good for confirming a paper exists and for seeing whether it is actually cited by others, which a fabricated paper never is.
If a reference is a preprint, especially in physics, mathematics, computer science, statistics, or economics, check arxiv.org directly. Every arXiv paper has a stable identifier and a permanent URL. A claimed arXiv ID that does not resolve to the stated title is fabricated. Note that arXiv is preprints, so a real arXiv paper may not be peer reviewed, which is a separate quality question from whether it exists.
Google Scholar is the broadest free academic index and catches gray literature, theses, and obscure venues the others miss. Use it to confirm existence when the specialist databases come up empty, but treat it as confirmation rather than proof, because it indexes more loosely and can surface a citation that only ever appeared in another AI-generated document. A title that appears only in Scholar, and nowhere in CrossRef, PubMed, or the publisher's own site, deserves a hard second look.
Here is the repeatable routine. It is built to be fast on the easy cases so you can spend your time on the suspicious ones.
Note where the references came from. A list pulled from default ChatGPT or Claude needs every entry checked. A Perplexity answer needs every link opened and read for relevance, but the existence check will mostly pass. This tells you where to spend effort.
For each reference that has a DOI, paste it through doi.org. This is the cheapest check and it clears most real references in seconds. Dead DOI means stop and investigate. Resolves to a different paper means the citation is a chimera. Resolves to the stated paper means move on. References with no DOI go to Step 3.
Take the exact title and search the database that fits the field: PubMed for medicine, Semantic Scholar or arXiv for CS and quantitative work, CrossRef or OpenAlex for everything else. Search the title in quotes. A real paper returns an exact or near-exact match with the same authors and year. No match in the field-appropriate database, after a title search, is your fabrication flag.
When you do get a hit, confirm the parts agree: the authors, the year, the journal or venue, and the volume or pages. Fabricated references often pair a real author with a journal they never published in, or shift the year. CrossRef and OpenAlex both show this metadata cleanly. If the author is real but the paper is not in their record, the citation is invented.
Existence is not enough. Open the actual paper and confirm it says what your draft claims it says. This is the step retrieval models do not save you from, and it is where most quiet errors hide. A real paper cited for a point it does not make is still a wrong citation.
Keep a simple checklist: verified, wrong-metadata, or fabricated. Replace anything that fails with a real source you have actually read, and never "fix" a fake citation by editing the DOI to one that resolves. That just buries the problem. If you found the fake through AI output, how to check if a ChatGPT citation is real covers the quick single-citation version of this same routine.
Realistically, the DOI resolve and title search clear most genuine references in under a minute each. The time goes into the handful that fail, and into Step 5, reading the source for relevance. For a 30-reference list, plan on an hour or two the first time and far less once the routine is muscle memory. That is a small price next to the cost of a fabricated reference reaching a reviewer, an editor, or, for lawyers, a judge.
Sooner or later a check fails. What you do in the next hour matters more than how the reference got there.
It is worth saying plainly, because the shame of it can make people freeze or, worse, hope nobody notices. AI models invent citations all the time, because of how they work, not because you did anything wrong. ChatGPT predicts plausible-sounding text one token at a time. It does not look anything up unless it is actively browsing, and even then it can misread what it finds. A reference is just a structured-looking string of text, so the model produces something with the right shape, an author, a year, a journal, a DOI, with none of it pointing anywhere real.
Researchers have measured how often this happens. Studies of AI-generated bibliographies have found that a meaningful share of citations are fully fabricated, with estimates commonly landing in the rough range of 15 to 30 percent, and a further chunk containing errors in the author, year, title, or page numbers even when the underlying paper is real. A widely cited 2023 analysis in Scientific Reports documented this pattern in detail. So if your model handed you a fake source, it was behaving exactly as the research predicts. You caught it, which is the system working.
That said, here is the part that does not change: you are responsible for what you submit. The model's mistake becomes yours the moment it goes into a paper, a brief, a report, or a post with your name on it. So we fix it properly rather than papering over it. The reassurance is real, and so is the responsibility.
The chapter on why a model invents sources explains the mechanism and the research behind it.
Before you tear anything out, make sure the source is actually fake and not just hard to find. Plenty of real papers are paywalled, badly indexed, or listed under a slightly different title. A quick check tells you which you are facing.
Work through these in order; most fakes fail the first two.
https://doi.org/ in your browser. A real DOI resolves to a real article page within a second or two. A fake one returns a "DOI Not Found" error from doi.org. This is the single fastest tell, and a dead DOI is strong evidence of invention.For the full step-by-step version, with the order to try things and what each result means, see how to check if a ChatGPT citation is real.
Not every fake is a clean invention. A sneaky variety is the chimera: a citation stitched together from real parts that never belonged together. The author is real, the journal is real, the year is plausible, but that author never published that title in that journal. These pass a lazy glance because every piece looks legitimate. To catch them, confirm the whole citation resolves as one unit to one real document, rather than checking each fragment separately. If the title search lands on a different paper, or the DOI resolves to a different title, you are looking at a chimera, and it counts as fabricated.
Here is the reframe that saves the most time. You do not need to rescue the fake citation. You need to support the claim it was attached to. The fabricated reference was the model's guess at evidence for something you wanted to say, so go find the genuine evidence for that statement. Pull the claim out of your sentence in plain words, then look for who actually established it.
Read enough of each candidate to be sure it actually says what you are claiming; a source merely about the same topic is not support for your specific point. This is where you do the scholarship the model skipped, and it is what makes the final work yours.
Once you have a real source that supports the point, swap it in cleanly and delete the fabricated reference entirely. Do not leave it in "just in case" and do not edit a fake DOI to look plausible, which only digs deeper.
Build the new citation from the real article itself, not from whatever the AI produced. Open the paper or its database record and read the author list, year, title, journal, volume, and pages straight from the source. Most databases and Google Scholar offer a "Cite" button that exports a formatted reference, but verify even those against the article, since automated exports occasionally garble a name or drop a page range. Then format it in the style your work requires: APA, MLA, Chicago, Bluebook, or a journal's house style.
Two small habits prevent the most common follow-on errors. First, make sure the citation in your text actually matches the claim it now supports; when you change the source, the sentence sometimes needs a small edit to reflect what the new paper found. Second, keep the link or DOI to the real article in your notes, so if anyone asks, you can resolve it on the spot.
This is the step people skip, and the one that protects you. If one citation was fabricated, treat every citation the model gave you as unverified until you have confirmed it. Fakes tend to cluster, because the conditions that produced one, the model guessing instead of retrieving, were in force for the whole session.
Make a checklist of every reference and run each through the same fast checks from Step 1: resolve the DOI, search the exact title, confirm the author wrote it. It goes quickly once you have a rhythm, and the peace of mind is worth far more than the ten or fifteen minutes. Pay closest attention to the citations that look most perfect, since the convincing ones are the chimeras that slip past a quick read.
If you are doing this for academic work, understand that a single fabricated citation can be treated as a serious matter even when it was unintentional, because the standard most institutions apply is about what you submitted, not what you meant. The chapter on misconduct explains how reviewers and integrity offices look at this, and why verifying every reference is the protection that holds up.
Once the fire is out, a couple of durable habits keep you from doing this again.
This is the core rule, and it resolves most of the problem on its own. AI is genuinely useful for the early, exploratory part of research. Ask it to explain a concept, suggest search terms, name the seminal thinkers, or outline a debate. Use those leads to find the real literature yourself in the databases above. What you must not do is paste an AI-generated reference into a finished document, because a model cannot reliably tell you a specific paper exists. Let it point you toward the library, not be the library.
Build verification into your workflow so it is not a special event you only remember after a scare. Before any piece of writing leaves your hands, run a final pass where you resolve every DOI and confirm every title against a real record. A simple "verified" column in your reference list or note-taking app makes this a habit, not a chore. The manual method is boring and reliable, exactly what you want between a model's guess and your name on a document.
Strip it down and the rescue is five moves. Confirm the source is genuinely fake with a DOI resolution and a title search. Find a real source for the same point through Google Scholar, the subject databases, and your library. Replace the citation and build it from the real article. Check every other reference, because fakes travel in packs. Then change your habits so AI does discovery and you do the citing.
You did the right thing by catching this before it counted. Finish the job properly, and you will not only be in the clear, you will have a verification routine that makes you more trustworthy than writers who never check.
Everything so far has been about catching a fabricated reference before it goes anywhere. This chapter is about what happens when one does not get caught, in academic work.
Academic integrity rules tend to draw a line that surprises people: intent is not always required for a finding. Plenty of institutions define misconduct to include conduct that is reckless or negligent, not only conduct that is deliberate. The logic is that the academy runs on trust. When you cite a source, you are vouching that you read it, that it says what you claim, and that it exists. A reader or grader should be able to follow your reference to the same thing you found.
A fabricated citation breaks that chain at the most basic level. The source is not real, so nobody can check it, and your claim now rests on nothing, regardless of who created the fake. If ChatGPT generated a confident-looking reference and you pasted it in without checking, the citation is still a misrepresentation of the evidence base, and that is the kind of thing a panel can treat as academic misconduct. "The tool made it up" explains how it happened, but it does not undo the fact that your work points to something that does not exist.
This is why "I didn't know" is a softer defense than people expect. It can affect the penalty, and a sympathetic instructor may treat a first slip as a learning moment rather than a formal case. But the underlying duty to verify what you cite sits with the author. That duty did not change when AI tools arrived. If anything, AI made it more important, because the tools produce fakes that look ordinary.
The chapter on why a model invents sources explains the mechanism in plain terms.
It helps to be precise, because not every flawed citation is fabricated. Reviewers and integrity panels usually see three kinds.
All three can land you in trouble, and the chimera and distorted kinds catch careful students off guard, because a quick glance makes them look fine. A more detailed breakdown lives on our companion page about the types of fake AI citations, but for misconduct purposes the principle is simple: if a reader cannot verify the source as you presented it, you have a problem to fix before you submit.
Honest numbers matter, so here are attributed ranges rather than a single scary figure. Studies that asked chatbots for bibliographies and then checked them found that a meaningful share of AI-generated citations are fully fabricated, with another portion containing errors in the author, year, title, or DOI. One widely cited analysis published in Nature's Scientific Reports examined ChatGPT's bibliographic output and found substantial fabrication, and follow-up tests report rates that vary a lot by model, field, and how the prompt was framed. The takeaway is not a precise percentage. It is that the rate is high enough that you cannot assume any AI-provided reference is real without checking it.
Two things have lowered those rates somewhat. Newer models hallucinate less than the 2023 versions, and tools with live web search or retrieval can pull genuine sources rather than guessing. But "less often" is not "never," and retrieval tools still misattribute or distort. The safe assumption for anything you will put your name on stays the same: treat every AI-supplied citation as unverified until you confirm it yourself.
The encouraging part is that catching fake citations does not require special software. The methods are mostly the same ones a careful person would use to look up any reference, so you can run the same checks before anyone else does.
At the course level, the most common trigger is a citation that does not resolve. A grader who tries to find a referenced paper and cannot, or who notices a DOI that leads nowhere, will look harder. Other tells: references suspiciously perfect in format but citing obscure or non-existent journals, a bibliography full of very recent papers none of them can locate, or in-text claims that do not match a real cited source. Some instructors spot-check a few references on every paper. Many more do it the moment something feels off.
In formal publishing, reviewers are expected to follow up on references that support key claims, and editors increasingly run submissions through reference-checking steps. The standard manual method is to resolve the DOI at doi.org, search the title in a database, and confirm the authors and venue match. Tools that cross-check references against CrossRef, PubMed, OpenAlex, or Semantic Scholar are becoming part of editorial workflows, and a citation that appears in none of them is a red flag. Journals have retracted or rejected papers over fabricated references, and a pattern of them raises questions about the whole work.
None of these checks are exotic. They are steps you can take yourself in a few minutes per citation. That is the heart of the fairness point: because the verification method is straightforward and available to you, "I didn't check" is hard to justify when the consequence is misrepresented research. We walk through the exact steps on our how to check if a ChatGPT citation is real page.
Outcomes vary widely by institution, by severity, and by whether it looks like an honest mistake or a deliberate one. From lightest to heaviest, the realistic range looks like this.
Where a case lands depends heavily on context, and intent usually moves the penalty even when it does not change whether something counts as a finding. A student who self-reports a mistake and shows their search records is in a very different position from one who cannot explain where a reference came from. That difference is why the habits below are worth building.
This is the part that actually helps, and it is not complicated. The goal is to use AI for what it is good at while keeping responsibility for accuracy where it belongs, with you.
Make verification a non-negotiable step, not an afterthought. For each reference, do the quick manual check: resolve the DOI at doi.org, search the exact title in Google Scholar or your library database, and confirm the authors, journal, year, and page numbers all match a real record. If a DOI leads nowhere, no database has the title, or the details do not line up, treat the citation as unverified and replace it. CrossRef, PubMed, OpenAlex, and Semantic Scholar are free and cover most fields. The whole check takes a minute or two per reference once you have the rhythm.
This is the single most useful rule. AI tools are good at helping you find directions, summarize a topic, suggest search terms, or point you toward an author or a debate. They are not a citation database, and they should never be the source of a reference in your bibliography. Use them to discover where to look, then go to the real database, find the actual paper, read enough to confirm it says what you think, and cite that.
Your research trail is your strongest protection. Save the search results, keep the PDFs or library links, and use version history in Google Docs or tracked changes in Word so there is a timeline of how your bibliography came together. If a reference is ever questioned, a record showing you found it in a real database settles the matter faster than any explanation, and it separates an honest mistake from something that looks evasive.
Policies differ, and they are changing fast. Some programs allow AI assistance with disclosure, some restrict it to certain tasks, and some require a statement of what you used and how. Read the policy that applies to you and follow it. When in doubt, disclose and ask. Undisclosed use that the policy required is its own integrity issue, separate from any fabricated reference, and it is entirely avoidable.
Fake AI citations can count as academic misconduct, and unintentional fabrication is not a free pass, because the accuracy of your references is your responsibility. That sounds strict, but it points to a fair and manageable standard: verify what you cite, use AI to find rather than to fabricate, keep your records, and disclose where required. Do those four things and the question of whether a fake citation is misconduct becomes one you never have to face.
The academic version of this costs a grade or a retraction. The legal version has cost lawyers their money, their reputations, and in a few cases their licence to practise. It is the clearest evidence available that this problem is not theoretical.
The reference point is Mata v. Avianca, Inc., decided in the Southern District of New York in 2023. The plaintiff's lawyers, Steven Schwartz and Peter LoDuca of Levidow, Levidow & Oberman, submitted a brief opposing a motion to dismiss that cited several decisions supporting their position on the airline's liability. The decisions were not real.
Schwartz had used ChatGPT to research the brief, and the model invented the cases outright: fictitious airline-injury decisions, complete with fabricated quotations and internal citations to other cases that also did not exist. The output looked authoritative because that is what a language model is built to produce, plausible text in the shape of what you asked for.
When the airline's counsel could not locate the cases and flagged the problem, it got worse rather than better. Rather than withdraw the fabricated authorities, the attorneys stood by them, and at one point produced what were presented as copies of the decisions, also generated by the tool. Judge P. Kevin Castel sanctioned the lawyers and imposed a $5,000 fine in June 2023.
The detail worth holding onto is the compounding error. The original mistake, trusting an unverified tool, is the kind a careful correction can contain. Judge Castel's Rule 11 finding rested on subjective bad faith, including acts of conscious avoidance and false or misleading statements to the court. A central aggravating factor was the failure to withdraw the fake cases once challenged, after there was reason to know they were false. If you take one lesson from Mata, take that one: the moment a citation is questioned, confirm it or pull it. Do not double down on a fabricated cite.
It would be convenient to file Mata as an early, freakish mistake from before lawyers understood the tools. The record does not support that. Since 2023, courts have documented at least fifteen cases involving AI hallucinated citations in filings, and they keep arriving. A few illustrate the range:
The common thread is not the size of the firm, the area of law, or the model used. It is the absence of a verification step between the tool's output and the signature. The technology changed. The duty did not.
For background on why these tools generate fabricated authority so confidently, our explainer on why ChatGPT makes up citations walks through the mechanism. In short, a language model predicts likely-sounding text; it does not look anything up, and has no sense of whether a citation corresponds to a real document.
Understanding the failure mode helps you guard against it. A large language model generates text one token at a time, choosing what is statistically probable given everything before it. When you ask for supporting authority, it produces strings shaped like citations, party names, a reporter, a volume and page, a year, because that form is overwhelmingly common in its training data. It is completing a pattern.
What the model does not do is consult a database of decisions and retrieve a real one. So the citation it returns can be entirely invented, or a chimera: a real case name welded to the wrong reporter, or a genuine citation paired with a holding the case never reached. Quotations are especially treacherous, because a fabricated quote reads exactly like a real one, with no stylistic tell.
This is why "but it sounded right" is neither a defense nor a safeguard. Sounding right is the one thing these tools reliably do. Some legal-research products now connect models to actual case databases, which reduces but does not eliminate the risk, because a tool can still mischaracterize a real case or pull an inapposite one. The signing attorney still has to read and confirm the authority.
None of this is a gap in the rules. The duty to verify a citation predates generative AI by a century, and the rules apply to AI-assisted work unchanged.
Rule 11 and its state analogues. When a lawyer signs a filing, Federal Rule of Civil Procedure 11 (and the equivalent state rules) certifies that the legal contentions are warranted by existing law, after an inquiry reasonable under the circumstances. Citing a case that does not exist is, by definition, not a reasonable inquiry. The rule imposes a gatekeeping duty on the signer, and that duty cannot be satisfied by trusting a tool's output on faith. The Mata sanctions rested on it.
Candor toward the tribunal. Model Rule of Professional Conduct 3.3 requires candor to the court, including not making false statements of law and correcting ones already made. A fabricated citation is a false statement of law. Once you know or have reason to know an authority is fake, candor requires you to correct it, which again points to a central aggravating factor in Mata: the cases were defended after challenge, which the court treated as part of a broader pattern of bad faith.
Competence and supervision. Rule 1.1 competence now includes a reasonable understanding of the tools you use, and Rules 5.1 and 5.3 put supervisory lawyers on the hook for work product from associates, paralegals, and non-lawyer assistants, including AI tools. A partner cannot offload verification to a junior who offloaded it to a chatbot. The signature carries the responsibility up the chain.
The throughline is simple: you cannot delegate verification to a chatbot. A generative tool can draft, summarize, and suggest. It cannot certify that a case exists or holds what you say it holds. Only a human who has read the authority can do that, and the rules assume it is you.
Verification is not complicated. It is a discipline, not a skill, and takes minutes per citation. Do it for every authority, every time, including the ones you are confident about, because confidence is what the fabrications exploit.
Start by confirming the decision is real and that the citation points to it. Pull it on Westlaw or Lexis, or look it up in the cited reporter, and check that the case appears, the party names match, and the reporter, volume, page, court, and year all line up with the brief. A citation that returns nothing, or a different case, is a red flag to resolve first.
Where you can reach the source directly, go to the court's own records. PACER for federal filings, or the court's docket and published-opinions system, will show you the genuine document. A real decision has a docket, a date, and a court of record behind it; a fabricated one does not.
Existence is necessary but not sufficient. A real case can be cited for a proposition it never stands for. Open the decision and read enough to confirm it holds what the brief claims. Check that the court, the procedural posture, and the holding are what you rely on, and that the case has not been reversed, vacated, or superseded. Run it through the citator (KeyCite or Shepard's) to confirm it remains good law.
Fabricated and distorted quotes are common, so treat quotations as their own verification task. For each quote, find the language in the actual opinion and confirm it appears, word for word, on the page the pin cite claims. Check that it is not stitched together from separate passages or lifted out of context in a way that changes its meaning. A quote you cannot locate verbatim in the opinion does not go in the brief.
If any part of a brief was drafted with a generative tool, assume every citation is unverified until you have personally confirmed it under steps one through three. Build the check into your workflow the way you already build in proofreading and cite-checking. Many firms now require a signed verification step for AI-assisted filings; even without a formal policy, Rule 11 requires it.
A useful habit: keep a short record of how each citation was confirmed (pulled on Westlaw, read in full, quote verified against page X). If a citation is ever questioned, that record lets you respond immediately, which is what Mata shows the courts want to see.
The chapter on checking one citation properly gives the mechanics of confirming a single reference, step by step. The earlier chapters collect the full picture of how and why these tools fabricate references.
The tools will keep changing. Models will get better at sounding certain, retrieval features will close some of the gap, and every few months somebody will announce the problem is solved. Treat that the way you would treat any other claim in this book, which is to say check it.
What does not change is the shape of the obligation. A citation is a promise to your reader that a specific document exists and says a specific thing. You are the one making that promise, not the model, and a promise you have not checked is a guess wearing a jacket.
The method in this book is not clever. Resolve the identifier, find the record, read enough of the source to confirm it says what you claim. A minute a reference. That is the whole trick, and it is the reason nobody who does it ends up in a sanctions order.
If you found this useful, send it to one person who needs it. Somebody drafting a literature review, somebody filing next week, somebody who has started trusting a chatbot a little too far.
And if the problem you keep hitting is the prose rather than the references, there is a companion book, Sounds Like You, on writing with AI without losing your own voice.
There is a third, Flagged in Translation, for anyone whose English started life in another language and keeps getting flagged for it.
You are welcome to share this book with anyone who might find it useful, as long as you share it whole and unchanged and do not sell it. Librarians, supervisors, editors and law firms: please link or hand it out freely, no permission needed.
Copyright © 2026 Dipak Bhosale. Published by Lacewing Technologies, Navi Mumbai, India.
This book is for general information. It is not legal advice and it is not a substitute for your institution's or your regulator's rules on AI use and citation. Verify any reference against the source databases before you rely on it. TextSight is an AI-detection and writing-trust tool; no detector, ours included, can confirm that a specific citation is real. Only the databases can do that.
A hallucination check does not verify a citation, and this book is clear that nothing except the databases can. What it does is show you which passages of AI output read as invented, so you know where the fabricated references are likely hiding before you start resolving identifiers one by one.