AI detectors and writing records answer different questions.
One guesses about a finished text. The other shows the text being written. The published numbers, plainly.
An AI detector reads a finished piece of writing and estimates the probability that a machine produced it. A writing record, which is what Receipts makes, captures the writing as it happens, every keystroke, pause, rewrite and paste, and replays it. These are not competing versions of the same thing. One produces an accusation that cannot be interrogated. The other produces evidence a person can watch and judge.
What the published research says about detectors
- A Stanford-affiliated study published in the journal Patterns (2023) tested seven widely used detectors and found they wrongly flagged 61% of TOEFL essays written by non-native English speakers, while performing near perfectly on essays by native speakers. The bias is structural: detectors reward predictable phrasing, and careful second-language writing is often predictable.
- Turnitin has publicly acknowledged a roughly 4% false-positive rate at the sentence level for its AI writing detection, and cautions against relying on scores below certain thresholds.
- Vanderbilt University disabled Turnitin's AI detector in 2023, noting that even a 1% false-positive rate across its roughly 75,000 annual papers would wrongly implicate around 750 students a year. Dozens of institutions have since disabled or declined AI detection, a list that has kept growing through 2026.
Sources: Liang et al., "GPT detectors are biased against non-native English writers," Patterns (2023); Turnitin's published false-positive statements; Vanderbilt University's public guidance on disabling AI detection.
The honest comparison
| AI detector (Turnitin, GPTZero and similar) | Writing record (Receipts) | |
|---|---|---|
| Input | The finished text only | The writing as it happens, keystroke by keystroke |
| Output | A probability score | A replay plus measured signals, no score, no verdict |
| Who runs it | The accuser, after suspicion | The student, before any accusation exists |
| Can it be questioned? | Limited; the scoring model is not open to inspection | Yes; anyone can watch the replay and judge |
| Failure mode | False accusation of an innocent student | An unconvincing record, which simply proves nothing |
| Non-native writers | Elevated false-flag rates documented across the seven detectors tested in Patterns (2023) | Typing signals do not depend on English fluency |
What a writing record cannot do, stated plainly: it cannot prove a student did not read AI output on another screen and retype it. No tool can prove that, including every detector on the market. Retyping does leave a visible signature, a flat rhythm with little revision and no abandoned wording, and the record makes that visible, but the judgment always belongs to a person. Receipts produces no verdicts on principle, because a tool that convicts students is the failure it exists to correct.
Looking for a Turnitin alternative?
Be precise about what you want to replace. If you want a different guess about finished text, other detectors exist, and they share the same structural limits. If what you actually want is to stop wrongly accusing honest students while still taking AI misuse seriously, that is not a detector at all: it is process evidence. Turnitin itself now sells a writing-process product (Clarity) that records students inside Turnitin's own editor for the school's use. Receipts approaches the same idea from the student's side: the record belongs to the writer, works wherever they write, stays free for every student, and no server of ours ever holds the writing. Which side of that line you buy depends on whether the record is kept about students for the school, or kept by students for themselves.
Is GPTZero accurate?
GPTZero publishes a false-positive rate at or below 1%. It was one of the seven detectors in the Patterns (2023) study that wrongly flagged 61% of TOEFL essays by non-native English speakers, which is the documented bias every probability-based detector shares. GPTZero also offers a Google Docs replay feature built on Google's revision history, which records periodic checkpoints rather than individual keystrokes. A keystroke-level record captures the timing texture inside the writing itself, and produces no score at all, because a score is exactly the thing an accused student cannot argue with.
What about Grammarly Authorship?
Grammarly Authorship is free and widely installed, and it labels text as typed, pasted or AI-assisted, provided the extension was running before writing began. It is a reasonable provenance label. The differences that matter: Authorship categorises, while a keystroke record replays; Authorship reports are generated by Grammarly's cloud service, while Receipts keeps writing in the student's browser and their own Google Drive; and Grammarly simultaneously sells AI writing and rewriting tools, which complicates what its attestation means. If a student already runs Grammarly, the two can even be used together.
They can coexist
Schools do not have to choose. Some keep a detector as a screening signal and use writing records as the evidence layer that protects students from the detector's mistakes. What no school should do, and what a growing number of universities have concluded, is treat a probability score as proof on its own.
See a record for yourself
Open the app, press Load sample, then open the Replay tab. Two minutes, no account.
See a live record School licences
Receipts — proof you wrote it. Operated by Receiptsproof Pte. Ltd. Home · Accused of AI? · For educators · For schools