Worth clearing up before anything else: Brisk Teaching does not produce an AI probability score. It is not a weak detector or a detector with a different threshold. It does a categorically different thing. Its writing analysis feature, Inspect Writing, replays how a document was composed: the keystrokes, the pastes, the deletions, the order paragraphs arrived in. Brisk's own description is about analysing submissions "for depth of thought, structural integrity, and originality", and understanding "how students revised their work". No percentage, no classifier verdict.
That is a genuinely good design decision, and we are not going to pretend otherwise to sell a comparison. If a student pastes 500 words in one event, a replay shows that happening, which is direct evidence about the writing process rather than a statistical guess about the finished text. Where it runs out is everything outside that window: work not composed in a connected editor, a document whose history is gone, or text you were simply handed. That is where a text signal is the only instrument left, and it is the honest split this page is about.
Read from its own site rather than from review pages, most of which on this keyword are written by services selling the opposite thing.
Brisk is a browser extension that works inside the tools teachers already use, with its best-known integration in Google Docs. Inspect Writing reconstructs a document's edit history and plays it back, so an instructor can watch the piece being built: typing, revising, reordering, and large paste events appearing as single jumps. Brisk describes offering two ways to understand how students revised their work, and frames the feature around depth of thought, structural integrity and originality rather than around AI classification.
The rest of Brisk is generative and feedback tooling: creating resources, giving written feedback, adjusting reading level. Brisk states it is used by more than 2 million teachers and it raised a Series A led by Bessemer. The extension is distributed through the Chrome Web Store. Plan pricing is not stated on the pages we were able to read, so we are not quoting a figure.
Process replay sidesteps the problem that makes detectors contentious. A detector infers authorship from the statistical properties of finished text, and it is wrong a measurable share of the time in a way that falls unevenly: Liang et al. (2023) found seven detectors flagging more than 61% of TOEFL essays by non-native English speakers on average, against near-zero for native-English US eighth-graders. A replay does not infer anything. It shows what the editor recorded.
Anthology reached a related conclusion from the vendor side, testing a market-leading AI detector for Blackboard, finding the error rate too high and the models biased, and publishing its decision not to build one. Brisk building process analysis instead of a classifier sits in the same tradition, and it deserves credit rather than a comparison table rigged against it.
Sources: briskteaching.com product pages and Chrome Web Store listing, read 29 September 2026; Liang et al., Patterns 2023. briskteaching.com/pricing returned 404 at the time of checking, so no price is quoted here.
They fail in opposite directions, which is the useful thing to understand and the reason this is not a ranking.
| Process replay (Brisk) | Text signal (a detector) | |
|---|---|---|
| What it observes | What the editor recorded | Statistical properties of the finished text |
| Certainty | Direct, not inferred | A probability, always |
| Needs a connected editor | Yes | No |
| Works on pasted or handed-over text | No | Yes |
| Works retrospectively, history gone | No | Yes |
| Bias against second-language English | Not applicable | Documented and significant |
| Defeated by drafting elsewhere and pasting once | Largely, yes | No |
The blind spots are structural rather than fixable. Work composed outside a connected editor and pasted in once leaves a single large paste event, which is ambiguous: it is equally consistent with drafting in Word offline, writing on a phone, and pasting from a model. Handwritten-then-typed work looks identical to generated work. Documents whose version history has aged out, or which were never in a supported tool, offer nothing to replay. And anything you receive as a finished file, which is the normal case for an editor, a client or a journal, has no process attached at all.
It is inference, and it is wrong a measurable share of the time. It carries a documented bias against second-language English and against formal register generally, because careful conventional prose is low-perplexity in the same way generated text is. It says nothing about who wrote the passage it flags. And it is degraded by rewriting: the largest academic comparison, Weber-Wulff et al. (2023) across 14 tools and 756 tests, measured average accuracy falling to about 26% on paraphrased AI text with about 71% going undetected.
They are complementary, and used together they are considerably better than either alone. A large paste event and a passage that reads as machine-written is a reason to have a conversation. A large paste event with unremarkable text is probably someone who drafts in Word. Clean process history with a high detector score is most likely a false positive, and that combination is the single most valuable thing either tool can tell you, because it stops a case before it starts.
These are the situations Brisk is not built for, and they are common.
This is the use Brisk is not aimed at at all, since it is a teacher-facing tool. If you are a student or a writer who wants to know which of your own sentences read as formulaic before anyone else looks, you need a text signal and sentence-level output. Our AI detector gives both, on 3 checks a day with no account. In our experience the passages that flag on genuinely human work are the introduction restating the question and the conclusion summarising the body, which is useful editing information regardless of anyone's policy.
If the accusation came from a detector, a second detector disagreeing is context rather than proof, and we would rather say that than sell you a certificate. Process evidence is the stronger card, and it is yours: version history is exactly the material Brisk-style replay reads, and you can produce it yourself from Google Docs or Word without any special tool. How to prove you did not use AI covers what to preserve and how to present it, and the appeal letter template is built around published research rather than around a counter-score.
Since the honest comparison above concedes real ground to Brisk, here is the same standard applied to us.
That our score proves anything. It does not, and for a teacher that matters more than any accuracy figure: a detector output is a reason to look at the work and talk to the student, never a finding. The institutions that handled this best either disabled detection or built a process where the score only ever triggers a conversation. Vanderbilt disabled Turnitin's AI detection in August 2023 and published why; the arithmetic it faced, roughly 75,000 papers a year, means even a 1% false positive rate implies about 750 wrongly flagged papers.
We also will not claim to be better than Brisk, because we do not do what Brisk does. If your students all write in connected Google Docs and you want to know how the work was made, process replay is the better instrument and we would rather you used it. Our own guidance for educators is at for teachers, and our limitations page lists where we fail.
Starting with the one that brings most people here.
Not in the sense of producing an AI probability score. Brisk's Inspect Writing replays how a document was composed, showing typing, revisions, deletions and paste events, and its own description is about analysing submissions for depth of thought, structural integrity and originality and understanding how students revised their work. That is direct evidence about the writing process rather than a statistical classification of the finished text, and it is a deliberate design choice rather than a limitation.
For what it covers, yes, because it observes what the editor recorded instead of inferring from statistics, and it carries none of the bias against second-language English that detectors do. It stops where the record stops: work drafted outside a connected editor and pasted in once, handwritten work later typed up, documents whose history has aged out, and anything you simply receive as a finished file. In those cases a text signal is the only instrument available.
Largely, by drafting elsewhere and pasting the result in a single event. That produces one large paste, which is ambiguous rather than damning: it is equally consistent with someone who drafts in Word offline, writes on a phone, or pastes from a model. This is exactly where the two approaches complement each other, since a large paste event combined with text that reads as machine-written is more informative than either observation on its own.
A file or email submission carries no replay, so process analysis has nothing to read and a text signal is what remains. Treat its output as a prompt to look at the work and talk to the student rather than as a finding. The most valuable combination is clean process history alongside a high detector score, because that pattern most likely indicates a false positive and stops a case before it starts.
Because of what they measure. Detectors read perplexity, how predictable your word choices are, and burstiness, how much that varies across a document. Careful second-language English tends toward a narrower, safer, more consistent vocabulary, which is statistically similar to generated text. Liang et al. (2023) measured seven detectors flagging more than 61% of TOEFL essays by non-native speakers on average, against near-zero for native-English US eighth-graders. Our detector carries the same bias and discloses it on the result.
No. We analyse text, not process, and we are not going to imply otherwise. We produce a statistical reading of finished writing with sentence-level highlights, accuracy published as bands of 88% to 92% on long-form English of 300 words or more and 70% to 78% under about 100 words, and a false-positive rate of 5.85% over 1,180 academic papers with the per-document data downloadable. If you want process evidence for work written in connected Google Docs, a replay tool is the right instrument.
That is the arrangement we would actually recommend, because the two fail in opposite directions. Process replay is direct but needs a connected editor and a surviving history. A text signal works on anything but is inference and carries a documented bias. Used together, a large paste event plus a formulaic-reading passage is a reason for a conversation, while clean history plus a high score is a reason to doubt the score.
How to use a detector output without turning it into an accusation.
Read the guidance →Anthology tested a detector, found the bias too high, and published its refusal to build one.
Read the detail →Where our own detector fails, listed by us rather than found by you.
Read the limits →The student side of process evidence, and how to produce version history yourself.
See the evidence list →1,180 academic papers, 5.85%, per-document data downloadable.
Check our numbers →What institutions actually run, and why the major platforms ship no detector.
Read the answer →For files, email submissions, contributor copy and anything retrospective, a text signal is what is left. Sentence-level highlights, accuracy published as bands, and a downloadable false-positive benchmark. 3 checks a day, no account.