Most do. The honest version of the answer is more useful than the yes: colleges very rarely buy an AI detector on purpose. They switch one on inside software they were already paying for. Turnitin added AI writing detection to a product that thousands of institutions already used for similarity checking, which is why the capability spread so fast and why so few students were told it had happened.
The scale is documented. Turnitin says that since launch in April 2023 it has reviewed over 200 million papers, of which over 22 million showed at least 20% AI writing. What is not documented is your own institution's policy, because that varies from department to department. Below: which tools actually get used, why your learning platform almost certainly has no detector of its own, what else gets looked at besides a score, and what the published error rates mean if a score lands on your work.
Five things that are true at almost every institution, and one thing that is true at none of them.
Assume your written coursework passes through some form of automated check. At the great majority of universities that check is Turnitin, and at many of those the AI writing indicator is enabled alongside the similarity score. You will often not see it. The AI indicator is frequently visible to the instructor and hidden from the student, which is the single most common reason people arrive at a page like this one confused about whether anything happened.
There is no single national or sector-wide answer. No regulator requires AI detection, no accreditation body mandates it, and there is no shared threshold at which a score becomes an allegation. Any page that tells you "colleges flag above 20%" as a general rule is inventing a standard that does not exist. The only authoritative source for your situation is your own institution's academic integrity policy and your module handbook.
The market is far more concentrated than the number of detectors on the internet suggests, and the reason is procurement rather than accuracy.
Turnitin is the dominant answer, and the mechanism matters more than the brand. Institutions bought Turnitin years ago for text-matching against published sources and a student paper archive. When AI writing detection was added to that existing product, every institution with a live licence acquired the capability without a new purchase, a new vendor review, or in many cases a new conversation with students.
Two consequences follow. First, adoption is much wider than deliberate demand for AI detection would predict. Second, Turnitin is not sold to individuals, so you cannot check your own work in the same tool your marker uses. That asymmetry, not accuracy, is the reason most people go looking for an alternative. Our Turnitin AI detector accuracy page holds the best-sourced version of its published error rates, and TextSight vs Turnitin sets out what is and is not comparable between the two.
Copyleaks and GPTZero both sell to institutions and appear in university guidance, GPTZero more often at individual instructor level because there is a usable free tier. Originality.ai shows up more in publishing and agency contexts than in universities. Beyond that you meet individual instructors pasting work into whatever free detector ranks first that week, which is the least controlled version of this and the one most likely to produce a bad outcome, because free consumer detectors are the ones least likely to publish a methodology or a threshold. We audit two of the most used in is GPTZero accurate and is ZeroGPT accurate.
Not every college that could check does. Vanderbilt University disabled Turnitin's AI detection in August 2023 and published its reasoning, which centred on not being able to see how the determination was made. The arithmetic behind that decision is worth keeping: Vanderbilt described grading in the region of 75,000 papers a year, so even a 1% false positive rate implies roughly 750 wrongly flagged papers annually. That is the calculation any honest institution has to do, and some of them have done it and walked away.
Sources: Turnitin press materials, April 2024; Vanderbilt University Center for Teaching statement, August 2023. Institution-level policies checked September 2026 and change frequently, so treat your own handbook as authoritative.
This is the most concrete public evidence of how much checking actually happens, and it comes from the vendor rather than from a survey.
In a release dated 9 April 2024, covering data to 21 March 2024, Turnitin reported the following about its AI writing detection feature, launched in April 2023:
| Turnitin's reported figure | Value |
|---|---|
| Papers reviewed since April 2023 | over 200 million |
| With at least 20% AI writing present | over 22 million (about 11%) |
| With at least 80% AI writing present | over 6 million (about 3%) |
Read that carefully, because it is routinely misread in both directions. It is not a measurement of how many students used AI, and it is not a measurement of detector accuracy. It is a count of how the tool scored the documents put in front of it, and it includes whatever false positives the tool produces. Turnitin is the party reporting both the numerator and the denominator here.
Two things, both useful. It settles the question this page is named after: checking at this volume is not hypothetical, and 200 million documents is not a pilot. And it gives you a realistic sense of base rates. If roughly 3% of submissions are scored at 80% or more AI writing, a high score is uncommon but not rare, which is precisely the range in which individual cases get handled badly.
The figures are also now more than two years old at the time of writing, and Turnitin has not published an equivalent update at the same level of detail. We have not found a newer figure from the company at this granularity, so this is reported as a 2024 snapshot rather than as a current rate.
Source: Turnitin press release dated 9 April 2024, reporting data to 21 March 2024. Figures are the vendor's own and are not independently audited. Captured 29 September 2026.
Moodle, Blackboard and Canvas are where students submit work, so people assume that is where detection happens. In all three cases it is not.
Moodle's own documentation states that it comes with no pre-installed plagiarism prevention methods. Everything depends on which third-party plugin an administrator chose to install, and the most-installed of those by a wide margin are Turnitin's. Moodle also has an AI subsystem, added in 4.5 and extended in 5.0, which generates, summarises and explains text. It contains no detection placement at all, which is worth knowing because "Moodle has AI features now, so it must detect AI" is a common and wrong inference. Full detail: does Moodle detect AI.
Blackboard ships SafeAssign, and SafeAssign is a similarity checker, not an AI detector. It compares submissions against internet sources, a ProQuest database, an institutional archive and a global reference database of volunteered papers. Anthology, which makes Blackboard, tested a market-leading AI detector in 2023, found the error rate too high and the models biased, and published its decision not to build AI detection. That is an unusually candid position from a vendor in this market. Full detail: does Blackboard detect AI.
Instructure built a socket rather than a detector. Canvas exposes a plagiarism platform through which an external tool registers itself, receives the submission by webhook and posts a report back. Canvas stores and displays that report. If your institution has not connected a tool to that socket, Canvas itself is checking nothing. Canvas quiz logs, often cited as evidence, record that a page lost focus and never what you opened in another window. Full detail: does Canvas detect AI.
It means the useful question is never "does my university use Canvas". It is "what is connected to it, and is the AI indicator switched on for this assignment". That is a question your module leader can answer in one sentence, and asking it before submission is a legitimate thing to do rather than a suspicious one.
Sources: docs.moodle.org, help.anthology.com, canvas.instructure.com plagiarism platform documentation, and marketplace.moodle.com plugin listings. Captured 28 September 2026.
In most real cases the detector score is not the evidence. It is the thing that made someone look, and what they look at next is far more informative.
This is the strongest signal on either side, and it belongs to you. A Google Docs or Word version history showing a document composed over days, with false starts, reordered paragraphs and abandoned sentences, is far better evidence of authorship than any percentage. A document that appears in two paste events is the opposite. Some tools now read this directly: Brisk Teaching, widely used in schools, replays a document's revision history rather than producing an AI score at all.
The practical consequence: keep your drafting history, and do not compose in a tool that discards it. Our guide on how to prove you did not use AI covers what to preserve and how to present it.
Markers who know your writing notice register shifts, vocabulary that appears from nowhere, and citation habits that change mid-module. This cuts both ways and is one reason genuine improvement can look suspicious, particularly for students who have just been through a writing support programme.
Where an institution has a decent process, the next step after a flag is a conversation: explain your argument, talk through why you structured it that way, discuss a source you cited. Someone who wrote a piece of work can almost always do this, and someone who did not usually cannot. If you are offered this, take it. It is the part of the process most likely to resolve in favour of an honest author.
Lockdown browsers and proctoring tools restrict or observe the exam environment. They do not read an essay and they produce no view about its authorship. A page that lists them as AI detection is padding.
Coursework and application essays run through different systems, on different policies, with different consequences, and they get conflated constantly.
Coursework flows through an institution's assessment stack, where a detector may already be enabled by default. An application essay goes to an admissions office, which is a separate function with separate software and, importantly, separate published policy. Some admissions systems screen for plagiarism. Whether AI detection is applied, and what weight it carries, is far less consistently disclosed than on the coursework side, and we are not going to assert a sector-wide practice we cannot source.
Read the specific instruction on the specific application. Many now state plainly what use of AI assistance is and is not permitted, and that stated rule is the one that governs, not a general assumption about detectors. Where a declaration is requested, answer it accurately. An accurate declaration of light assistance is a much better position than an inaccurate denial, and it removes the detector from the conversation entirely.
The one thing a personal statement has going for it is that it is personal. Specific, verifiable, particular detail about your own experience is both what admissions readers are looking for and, incidentally, the hardest thing for a model to produce convincingly. Writing it properly is the strategy.
If your institution checks, you carry a share of that tool's error rate whether or not you have ever used a language model. This is the part worth understanding properly.
Turnitin's own documentation puts its false positive rate at under 1% at document level above the 20% threshold, and about 4% across the full distribution, measured on an internal evaluation set whose size is not disclosed. It recommends not acting on scores below 20%. Those are the vendor's numbers for the vendor's tool, reported here as theirs.
Across the wider field the numbers are worse. The largest academic comparison to date, Weber-Wulff et al. (2023), tested 14 tools across 756 cases and concluded they "are neither accurate nor reliable." It measured a field-wide false positive rate averaging about 2% on human text, rising to about 11% once machine translation was involved. Turnitin scored highest of the 14.
If English is not your first language, the risk is not evenly distributed. Liang et al. (2023), published in Patterns, ran seven detectors over TOEFL essays written by non-native English speakers and found they flagged more than 61% of them on average, against near-zero false positives on essays by native-English US eighth-graders. The seven were Originality.AI, Quil.org, Sapling, OpenAI's classifier, Crossplag, GPTZero and ZeroGPT. No per-tool breakdown was published, so the figure describes the group and not any single product.
The mechanism is not mysterious. These detectors read perplexity, meaning how predictable your word choices are, and burstiness, meaning how much that predictability varies. Careful second-language English tends to use a narrower, safer, more consistent vocabulary, which is statistically indistinguishable from the thing the model is looking for. The classifier is reading the pattern correctly and drawing the wrong conclusion from it. More on this in AI detector false positives and why AI detectors get it wrong.
The same bias applies to formal academic register generally, to heavily edited prose, and to writing produced with grammar tools. If you have polished a draft hard, you have moved it toward the statistical profile of generated text.
Sources: Turnitin published documentation (2024); Weber-Wulff et al., International Journal for Educational Integrity 2023, arXiv:2306.15666; Liang et al., Patterns 2023. TextSight was not evaluated in either study and we do not construct a comparable figure for ourselves by analogy.
In order, and the first one is the one people skip.
Ask which tool produced the number, what the number was, and what threshold your institution applies. You cannot respond to "the system flagged it." A named tool, a figure and a policy threshold are the minimum you need, and you are entitled to ask for them.
Version history, drafts, notes, search history, library loans, the messy artefacts of having actually done the work. This is stronger than any counter-score, and it is the material an appeal should be built on. Retrieve it early: some platforms age version history out.
Disputing whether you are the 8% the tool got wrong is an argument you cannot win with assertion. The published literature is the lever: the documented false positive rates, the non-native English finding, the vendor's own guidance about thresholds and about not treating a score as decisive. Our AI detection appeal letter template is built around exactly that structure, and accused of using AI in college walks through the process end to end.
A second detector disagreeing is genuinely useful context, because it demonstrates that the tools are not measuring a fact about your document. It is context and not proof. Two detectors agreeing does not make either right, and a favourable score from us is not evidence of your innocence. We would rather say that plainly than sell you a number to wave at a committee.
You are reading this on a detector vendor's website, so here is the disclosure and the limits, in the same place.
Checking your own writing before you submit it, and seeing which sentences read as machine-written rather than receiving a single number. Our AI detector shows sentence-level highlights, which is the part that tells you something actionable: usually it is the introduction restating the question and the conclusion summarising it, the two places where careful human writing is most formulaic. The free tier runs 3 checks a day with no account.
If you want the measurement detail rather than the summary, our accuracy methodology sets out how those bands were produced and on what data.
Answered from published sources, with the honest "it depends on your institution" where that is genuinely the answer.
No, and there is no requirement that they do. Checking is common rather than universal, it is usually enabled inside plagiarism software the institution already licenses, and several well-known universities have deliberately switched it off. Vanderbilt disabled Turnitin's AI detection in August 2023 and published its reasoning. Because policy is set institution by institution and sometimes department by department, your academic integrity policy and module handbook are the only authoritative answer for your situation.
Turnitin, by a wide margin, and mostly because it was already installed for similarity checking before AI detection was added to it. Copyleaks and GPTZero also sell into education, and individual instructors sometimes use free consumer detectors on their own initiative, which is the least controlled version of this. Turnitin is not sold to individuals, so you cannot check your work in the same tool your marker uses.
None of the three ships an AI detector of its own. Moodle's documentation states it comes with no pre-installed plagiarism prevention methods, so everything depends on which plugin an administrator installed. Blackboard ships SafeAssign, which is a similarity checker, and Anthology published its decision not to build AI detection after finding error rates too high and the models biased. Canvas provides a plagiarism platform that an external tool plugs into, and detects nothing itself if nothing is connected.
Turnitin reported that of over 200 million papers reviewed between April 2023 and March 2024, over 22 million (about 11%) showed at least 20% AI writing present and over 6 million (about 3%) showed at least 80%. Those are the vendor's own unaudited figures, they describe how the tool scored documents rather than how many students used AI, and they include whatever false positives the tool produces. No comparably detailed update has been published since.
A score on its own is not proof, and the better-run institutions say so in their own guidance. Turnitin recommends not acting on scores below its 20% threshold, and Grammarly states that no AI detector can conclusively determine whether AI was used. In practice a score is what prompts someone to look, and what follows is where cases are actually decided: version history, drafts, consistency with your other work, and a conversation about your argument. Our appeal letter template is built to move the discussion onto that ground.
The risk is measurably higher, and this is the most important thing on this page for anyone in that position. Liang et al. (2023) ran seven detectors over TOEFL essays by non-native English speakers and found they flagged more than 61% on average, against near-zero for native-English US eighth-graders. Careful second-language English uses a narrower, more consistent vocabulary, which is statistically similar to what these tools look for. Keep your drafting history, and if you are flagged, cite that research rather than only asserting authorship.
It is worth doing, as long as you know what the result is worth. A flag on your own genuine writing tells you which passages read as formulaic, which is useful editing information and usually points at your introduction and conclusion. It does not predict what your institution's tool will say, because different detectors disagree constantly, and a clean result from us is not a guarantee about a different model. Use it to improve the draft, not to certify it.
That is a separate system from coursework and far less consistently disclosed, so we will not assert a sector-wide practice. Read the specific instruction on the specific application, because many now state plainly what AI assistance is permitted, and that stated rule governs rather than any assumption about detectors. Where a declaration is requested, answer it accurately: an accurate declaration of light assistance is a far better position than an inaccurate denial.
The process end to end, from the first email to the committee, written for the person it is happening to.
Read the guide →A template built around published research rather than around asserting your innocence.
Use the template →What process evidence to keep, and how to present version history so it carries weight.
See the evidence list →The socket, not the detector. What Canvas actually does with a submission, and what quiz logs record.
Read the detail →Its published false positive rates, the 20% threshold, and what the figures do and do not cover.
Check the numbers →Why sub-1% claims collapse on real writing, and which groups carry most of the risk.
Read the guide →See which sentences read as machine-written, at sentence level, with the false-positive benchmark published beside the score. 3 checks a day, no account, no card. We will not tell you a result makes you safe, because it does not.