ResearchBenchmarkFalse PositivesAI DetectionNon-Native EnglishAccuracy

How Often Do AI Detectors Flag Human Writing? We Published Our False-Positive Benchmark

We tested 715 passages with five-fold cross-validation: a 4.9% false-positive rate on human writing and 2.9% on non-native English. Our published benchmark.

Paul Byrne··5 min read


Most AI detectors will tell you how much AI they catch. Almost none will tell you how often they are wrong about a real person. That second number is the one that matters, because a false positive is not a statistic. It is a student accused of cheating for work they wrote themselves.

So we are publishing ours.

The real harm is the false positive

When an AI detector flags human writing, someone can lose marks, face a misconduct hearing, or be told to rewrite work that was already their own. The cost of a false accusation is far higher than the cost of missing one piece of AI text. Yet the industry markets on catch rate, not on how rarely it gets an innocent person wrong.

There is a specific fairness problem underneath this. A 2023 Stanford study (Liang et al.) found that GPT detectors systematically flagged writing by non-native English speakers as AI-generated, because non-native and formal academic writing shares surface patterns with model output. The people most likely to be wrongly accused are the ones least equipped to contest it. We have since taken this a step further on our own corpus and published which kinds of writing our detector falsely flags, split by dialect and register: British English came out at zero of 85, while Wikipedia-style encyclopaedic prose carried 81% of the false positives.

The benchmark: what we measured

We ran our detector across 715 passages: 391 AI-generated (from Claude, GPT and Gemini), 324 written by humans, of which 35 were written by non-native English speakers. We used five-fold stratified cross-validation, so every score is out-of-sample. Here is the shipped model.

MetricResult

AI catch rate (recall)54.7%
False-positive rate on human writing4.9%
False-positive rate on non-native English2.9%
Precision93.0%

Two honest things to sit with. First, a 4.9% false-positive rate means roughly one human passage in twenty gets flagged. That is not zero, and no one should treat a single flag as proof. Second, we catch about 55% of AI text, which means we miss close to half. Detection is genuinely hard, and any tool claiming near-certainty is overselling.

The number we are proudest of is the non-native one: 2.9%, lower than our overall human rate, not higher. On the cohort the Stanford research showed most at risk, we do not over-flag.

The trade we made on purpose

While tuning the model, we built a version that caught 62.4% of AI text, nearly eight points more than what we shipped. We did not ship it. That version flagged 14.3% of non-native English writers as AI, a fivefold jump in exactly the group least able to defend themselves.

We chose the model that catches less AI and protects non-native writers. Given the choice between catching more cheats and falsely accusing more real students, we think the answer is obvious, and we would rather be honest about making it than quietly optimise for a bigger headline number.

What this means if you are checking work

  • Treat any result as a screening signal, never as proof. A 4.9% false-positive rate is low, but on a class of 200 essays it still means flags you must not act on blindly.

  • Read the flagged passages, not just the score. The value is in seeing which sentences triggered the model, then judging them in context.

  • Be especially careful with non-native English writers. Even a fair detector is a starting point for a conversation, not evidence for a hearing.

Can any AI detector be 99% accurate?

You will see 99% accuracy claims across this industry. Treat them as marketing until you can answer three questions about the benchmark behind them. What was in the test set, and did it include edited, paraphrased and mixed text rather than raw model output? What was the false-positive rate on human writing, reported separately rather than blended into one headline figure? And was the test run on writing the model had never seen?

There is no official certification a detector can pass. The closest thing to a neutral standard is NIST's GenAI Text Challenge, which ran a pilot text discrimination evaluation in 2024 and has an open discriminator track for 2026. It scores detectors on discriminative power and calibration, not on a single accuracy number, and no commercial detector's 99% marketing claim comes from it.

Our own benchmark below is the standard we think every vendor should meet: out-of-sample results, false positives reported separately, and the failure cases included.

Benchmark methodology

The 391 AI passages span Claude (Haiku and Sonnet), GPT and Gemini, including passages post-processed to read as tired student writing. Human passages include formal academic prose and the non-native English cohort. Scores come from a stylometric logistic-regression model with a fingerprint booster, cross-validated on a fixed seed, with verdict bands published in our methodology notes. We report out-of-sample numbers only, because in-sample accuracy tells you nothing about how a detector behaves on writing it has never seen.

We will update these figures as the model changes. If you want to check a piece of writing against this detector, you can run three free scans a day at isitai.co.uk with no account.

Frequently asked questions

How often do AI detectors flag human writing as AI?

It depends on the detector and the writer. When we tested our own detector on 715 passages we measured a 4.9 percent false-positive rate on human writing overall, and 2.9 percent on non-native English writing. Independent testing of other detectors has found false-positive rates far higher, up to around 61 percent on non-native English text, which is why a single detector flag should never be treated as proof.

What is a good false-positive rate for an AI detector?

Lower is better, because a false positive means a real human is wrongly accused. There is no official standard, but a rate in the low single digits on ordinary human writing is a reasonable bar, and the rate on non-native English writing matters most because that group is flagged disproportionately. We publish ours openly: 4.9 percent overall and 2.9 percent on non-native English.

Why do AI detectors flag non-native English writers more often?

AI-generated text and non-native English writing can share surface features: simpler sentence structures, more predictable word choices, and less idiomatic phrasing. Detectors that chase a high catch rate tend to over-flag these patterns, which penalises non-native writers unfairly. We deliberately capped our false-positive rate rather than maximise catch rate, accepting a lower recall of 54.7 percent to protect human writers from wrong accusations.

Should a teacher rely on an AI detector flag as proof?

No. Even a low false-positive rate means some human work will be flagged, and the rate is higher for non-native English writers. A detector result is a screening signal, not evidence. It should start a conversation, prompt a look at draft history and process, and be weighed alongside what the teacher already knows about the student, never used as proof on its own.

Can an AI detector be 99% accurate?

Not in any sense that matters in a classroom. Vendor 99% claims come from internal benchmarks that pit raw AI text against raw human text and exclude the hard cases: edited output, paraphrased text, mixed essays, formal academic prose and non-native English writing. Our own out-of-sample benchmark catches 54.7 percent of AI text with a 4.9 percent false-positive rate, and we publish both numbers because a single accuracy headline hides the trade-off between catching AI and wrongly flagging humans.

Is there an official benchmark for AI text detectors?

There is no official certification a commercial detector can pass. The closest neutral standard is the NIST GenAI Text Challenge, which ran a pilot text discrimination evaluation in 2024 and has an open discriminator track for 2026. It scores detectors on discriminative power (AUC) and calibration (Brier scores) rather than a single accuracy figure, and no commercial detector's marketing claim is based on it. Until a shared public benchmark exists, ask any vendor for out-of-sample results with false positives reported separately.

Try Is It AI?

Detect AI-generated content instantly. 3 free scans per day.

Scan Content Now

Free AI text check

Free, no signup

Try Now