aiwut?AI tools, decoded

Education

wut is GPTZero?

The AI detector teachers reach for when the school has not bought one.

Wut’s the catch

The gap between the marketing and the measurements. Vendor accuracy claims in this category run 87–99%, while independent testing has put GPTZero's false-positive rate as high as 12% in one comparison and around 23% on real student essays in classroom testing — one analysis estimated a teacher relying on it would falsely accuse roughly one in five innocent students.

How it scores

Transparency
Can you find out what it costs, who owns it, and how it works — before you pay?
Privacy
What happens to what you feed it, and who else gets to see it.
Value
Does the thing it charges for actually work well enough to be worth the money?
Staying power
Will this still exist, and still be the same product, in two years?

Wut it actually does

You paste text and GPTZero returns a probability that it was machine-generated, with a sentence-level highlight of the passages it considers most suspicious. It is fast, free at low volume, and requires no institutional purchase — which is exactly why individual teachers use it.

That accessibility is the whole product and the whole problem. Turnitin at least arrives wrapped in institutional policy and, sometimes, an appeals process. A detector a teacher found in a browser tab arrives with neither.

The measured error rates

Independent testing has produced consistently worse numbers than vendor claims. One 2026 comparison of five detectors put GPTZero's false-positive rate at 12%, the highest in that test. Independent classroom testing on student essays found around 23%. Analysis by Futurism estimated that a teacher relying on GPTZero would falsely accuse roughly 20% of innocent students.

Set that beside category accuracy claims of 87–99% and the distance is the story. These figures vary because they are measuring different corpora — that variation is real and worth acknowledging rather than cherry-picking the worst — but every independent number is far from "reliable enough to accuse someone."

The same Stanford work covering the whole category found detectors misclassifying non-native English writing at an average 61.3% false-positive rate. Whatever the exact figure for any one tool, the direction is consistent and the burden falls in the same place.

Why a percentage is the wrong output

A detector that returns "89% likely AI" is producing something that looks like evidence and behaves like a hunch. There is no way for the teacher to interrogate it, no way for the student to rebut it, and no external record of how often that particular number has been wrong.

GPTZero does publish compliance credentials — SOC 2 Type II, and adherence to FERPA guidance on student data. That is real and worth crediting: it means the tool handles student data responsibly. It says nothing whatsoever about whether the score is right, and those two things get conflated constantly in procurement.

The honest framing is that these tools estimate how typical a piece of writing is. Unusually typical writing gets flagged. Second-language writers, students taught to a rigid structure, and anyone who writes carefully and plainly all produce unusually typical writing.

The verdict

Accuracy is the wrong number to argue about. What matters is that a false positive here is an accusation of academic dishonesty against a specific teenager, made by a teacher who has been handed a confident-looking percentage and no meaningful way to check it. Nothing at these error rates should be used to decide whether a person cheated.

Reasonable as a private prompt to look more closely at a piece of work. Not reasonable as evidence, and not reasonable at all where the student cannot see the score and respond. If you are the one accused, the independent numbers below are the most useful thing on this page.

This section is our opinion. Everything stated as fact above is sourced below.

More in AI detectors

Sources

Facts last checked against these on 2026-07-30.

  1. ProofreaderPro — how accurate are AI detectors in 2026, five testedGPTZero showing the highest false-positive rate (12%) in that five-detector comparison, against category accuracy claims of 87–99%.
  2. Working Educators — GPTZero review 2026: AI detection accuracy for teachersIndependent classroom testing finding ~23% false positives on student essays; the Futurism estimate of ~20% of innocent students falsely accused; SOC 2 Type II and FERPA claims.
  3. Stanford / Patterns — GPT detectors are biased against non-native English writersThe 61.3% average false-positive rate on non-native English writing across seven detectors.
  4. GPTZero privacy policyData handling for submitted text and student data.