If you have been flagged and you need to prove you didn't use AI, the honest answer is that no single thing settles it, and the strongest evidence is not a detector score at all. It is the record of how you wrote the piece: the timestamps, the messy drafts, the half-finished sentences you later cut. This guide ranks the evidence by how convincing it actually is and shows you how to gather and present each piece so a person reviewing your case can see, plainly, that the work is yours.
This is the evidence methods reference for our falsely accused of using AI hub. If you want the wording of the message you send, see our AI detection appeal letter guide. If you want to understand the formal college process, your rights, and the timeline, read accused of using AI in college. This page is about the proof itself.
The first thing to understand is that you are being asked to prove a negative, which is hard for anyone. A detector flagged your work as likely machine-generated. That flag is a probability, not a fact, and the people on the other side of the table know that, even when they act like they don't.
So the goal is not to find one magic document that ends the conversation. The goal is to build a picture that only a real author could have. A finished essay tells almost nothing, because both you and a chatbot can produce a clean final draft. What a chatbot cannot produce is your specific, dated, evolving process: the paragraph you wrote in the wrong order, the source you opened at 11 p.m. on a Tuesday, the sentence you rewrote four times. Evidence of process is the thing. Everything below is ranked by how directly it shows process.
If you wrote in Google Docs or Microsoft Word, you may already be holding the single most convincing piece of evidence without realizing it. Both tools quietly record the writing as it happened.
Google Docs Version History. Open the document, go to File, then Version history, then See version history. You will see a timeline of every meaningful save, with timestamps, going back to the blank page. This shows the document growing: a few sentences, then a paragraph, then a rearrangement, then edits over days. That gradual build is almost impossible to fake and is exactly what AI-pasted work lacks, because pasted text appears all at once in a single jump.
Word Track Changes and AutoRecover. If you used desktop Word, the document's editing history, autosave versions, and file metadata (created date, modified date, total editing time) tell a similar story. If the file lives in OneDrive or SharePoint, there is a full version history there too, similar to Google Docs.
Do not just say "I have version history." Show it. Take dated screenshots of the timeline, and if you can, share the live document (view access) or export the version history so the reviewer can scrub through it themselves. Point to specific moments: "Here, on the 3rd, the conclusion didn't exist yet. Here, on the 5th, I rewrote it." A reviewer watching a document grow over real time, in your account, is the closest thing to a confession of innocence that exists. This is why we put it first.
One caution, stated honestly: version history is powerful only if it exists. If you drafted in a plain text editor, or copied your own writing in from your phone's notes in one paste, the timeline can look thinner than your actual effort. That does not mean you cheated. It means you may need to lean harder on the other evidence below.
The next tier is anything that captures your thinking before the final version existed.
Gather these into one folder, in rough date order, and write a one-line caption for each: what it is and when it was made. The story you are telling is "idea, to outline, to rough draft, to final," and each artifact is a frame in that story. A reviewer should be able to follow the trail without you narrating every step.
Your research footprint corroborates the rest. Browser history, your library's database access logs, open tabs, downloaded PDFs, and saved search results all show you actually engaging with the sources you cite. If your bibliography matches a trail of real visits to those exact pages and articles, that alignment is meaningful. It does not prove you wrote every sentence, but it proves you did the reading the essay is built on, which a chatbot does not do on your account, at your times.
Present this as a supplement, not a centerpiece. A short list of "sources I accessed and when, matching my citations" placed alongside your drafts strengthens the whole package.
Here is where a detector fits, and it is important to be precise about what it can and cannot do.
A detector cannot prove you didn't use AI. No detector can, because it returns a probability, not a fact, and it can be wrong in either direction. So a low AI score is not a verdict in your favor. What it is, used carefully, is one additional data point: a second, independent opinion that disagrees with the first.
That disagreement is the useful part. If the tool that flagged you said "likely AI" and a different, independent tool says "likely human," you can fairly point out that two detectors looking at the same text reached opposite conclusions, which is exactly what you would expect from probabilistic tools with real error rates. That variance undercuts the idea that any single score is reliable enough to act on. You are not claiming the second tool is right. You are showing that detectors disagree, so no one score should decide your case.
You can run your own text through TextSight's AI Detector to get that second score plus a sentence-by-sentence breakdown, which is more useful than a single number because it shows you (and a reviewer) which specific lines triggered the original flag. Frame it honestly when you present it: "An independent detector scored this differently and flagged different sentences, which shows these tools are not consistent." Never present it as proof. It is a supporting data point that lives below your version history and drafts in the ranking, not above them.
The most disarming thing you can do is volunteer to be questioned. Offer to sit down with your instructor and talk through your essay: explain your argument, defend a particular choice, recall why you cut a section, summarize a source in your own words. Someone who genuinely wrote a piece can do this easily, and the offer alone signals confidence. A short note like "I'm happy to walk through my drafting process or answer questions about any part of this in person" reframes you from accused to cooperative, and it puts the burden back on a single tool to justify itself against a real conversation.
Put everything in one place, ordered by strength, with a short index at the top. A clean package does half the persuading for you. A workable structure:
Keep the tone calm and factual. You are not pleading. You are presenting a record. For the exact email wording that carries this pack, use the AI detection appeal letter template, and route it through your academic integrity office or instructor following the steps in accused of using AI in college, not by arguing with the detector vendor.
It helps to know, and to be ready to cite, why these flags are shaky in the first place. A Stanford study (Liang et al., 2023) found that seven popular GPT detectors flagged essays by non-native English writers as AI at an average false-positive rate of 61.3 percent, with one tool hitting 97.8 percent, roughly six times the rate for native writers. Independent estimates of general false-positive rates run around 5 to 20 percent. The University of Texas found that about 11 percent of papers flagged on its campus were false positives on review. And Turnitin itself states that its AI score "should not be used as the sole basis for adverse actions against a student." Several universities, Vanderbilt among the first, disabled AI detection over these reliability concerns.
You do not need to attack any tool. You just need to put these attributed facts next to your evidence so a fair reviewer sees that the flag is one uncertain signal and your record is many concrete ones.
TextSight is an AI-detection and writing-trust tool. We help you check and improve your own work and gather a second, independent opinion. No detector, ours included, can prove that you did or did not use AI, because every score is a probability and can be wrong. The strongest evidence is always your writing process: version history, drafts, and notes. Use a detector as one supporting data point, never as proof.
TextSight's free tier gives you daily scans with sentence-level highlights, so you can see which lines carry the AI signal and why. Every score is a probability, not proof, and we say so plainly.