AccuracyFalse PositivesBritish EnglishMethodologyData

Does British English Trigger AI Detectors? We Tested Ours on 324 Human Passages

We split our 324-passage human calibration corpus by dialect and register and re-ran our shipped detector. British English was flagged 0 times in 85 passages. The writing that actually gets falsely accused is something else entirely, and we publish the full numbers.

Paul Byrne··6 min read


A question keeps appearing in the retrieval logs that AI engines generate when people ask them about detectors: how accurate are AI content detectors on British English? It is a fair worry. Most detection tools are built and calibrated in the United States, on American training data, and a teacher in Leeds might reasonably wonder whether "colour", "whilst" and long Guardian-style sentences read as anomalies to a model that has mostly seen American prose.

Nobody in this category publishes an answer with data attached. So here is ours, from our own calibration corpus, with the numbers that make us look good and the numbers that do not.

What we tested

Our calibration corpus contains 324 human-written passages alongside 391 AI-generated ones. The human side spans Guardian journalism and opinion columns, classic literature from Project Gutenberg, Wikipedia articles, academic paper abstracts, forum and discussion posts, and writing by non-native English speakers. Every passage carries source metadata, which is what makes a dialect split possible.

We grouped the human passages by source. Guardian journalism is British English by construction. Project Gutenberg authors were classified by nationality where it is unambiguous: Dickens, Austen, the Brontes, Conan Doyle, Stevenson and company on the British side; Twain, Thoreau, Melville, Du Bois and company on the American side. Authors who wrote in translation, or whose text is centuries old, sat in their own group. Wikipedia, academic and forum sources stayed as their own registers.

Then we scored every passage with the same detector configuration that runs on this site, using out-of-sample predictions: the model never sees the passage it is scoring during training. This is the configuration behind the accuracy figures we already publish, which came out at 4.9 per cent false positives overall, with 54.7 per cent of AI passages caught.

The result: British English was flagged zero times

Across 85 passages of identifiably British writing, 40 pieces of Guardian journalism and 45 passages of British classic literature, the detector falsely flagged none. Zero of 85. With a sample that size, statistics obliges us to say the true rate could be as high as 4.3 per cent, but the observed rate was zero, and it was also zero on the American classics we tested, so this is not a case of the detector favouring one dialect over the other.

GroupPassagesFalsely flaggedRate

British journalism (Guardian)4000%
British classic literature4500%
American classic literature2900%
Academic writing4500%
Forum and discussion posts5012.0%
Non-native English writers3512.9%
Wikipedia articles801316.2%

If dialect were the problem, British spellings and British sentence rhythm would show up in that table. They do not.

What actually gets falsely accused

Look at the last row. Of the 16 false positives in the entire corpus, 13 were Wikipedia articles. That is 81 per cent of our false accusations landing on one kind of writing, at a rate of 16.2 per cent on that source, versus zero on journalism, zero on classic fiction and zero on academic abstracts.

This makes uncomfortable sense once you think about what statistical detection measures. Encyclopaedic writing is neutral in tone, uniform in sentence length, dense with facts, and stripped of personal voice, because those are Wikipedia's house rules. They are also, almost point for point, the properties detectors treat as signatures of AI text. AI models write the way they do partly because they were trained on text like Wikipedia. A human who writes in that register inherits the suspicion.

The two literary exceptions prove the same rule from the other direction. One flagged passage was from Dubliners, where Joyce's flat, deliberately unadorned narration reads as uniform to a statistical model. The other was Beowulf. Our detector falsely flagged a poem that predates the Norman conquest, which we offer as a permanent cure for anyone tempted to treat a detector verdict as proof.

The practical translation for teachers: the student essay most at risk of a false accusation is not the one written in any particular dialect. It is the carefully neutral, fact-dense, personality-free essay, the exact style some students are taught to produce for formal assessment. This is why we say a detector score is a starting point for a conversation and never evidence on its own, and why false positives are the number that matters most when you compare tools.

The number that does not flatter us

An honest study reports the abstentions. On British journalism our detector said "uncertain" 25 times out of 40, a 62 per cent abstention rate, far higher than on any other group. It never called a Guardian column AI, but it frequently declined to call it confidently human.

We think we know why: modern professional opinion writing is polished, structurally controlled and edited toward consistency, which pushes it toward the statistical middle ground where our thresholds refuse to commit. Refusing to commit is the behaviour we chose on purpose, because the alternative, forcing a verdict, is how detectors generate false accusations. But a 62 per cent uncertain rate on professional British journalism is a real limitation, we are publishing it because you should know it exists, and reducing it without raising the false positive rate is on our calibration roadmap.

Method notes and limits

Everything above comes from one corpus and one detector, ours. The sample sizes are honest but modest: 85 British passages bounds the true false positive rate below 4.3 per cent at 95 per cent confidence, not below zero. Dialect classification is by source and author nationality, not by spelling analysis of individual passages. Writers in translation and pre-modern texts resist dialect labels, which is why they sat in their own group. We have not tested other vendors' tools on this corpus, so this study says nothing about whether GPTZero or Turnitin handle British English well; it says only that ours does, and it shows the kind of evidence any vendor could publish if they chose to.

The full accuracy figures behind our shipped detector, including how often detectors flag human writing generally and how our accuracy compares with other tools' published claims, are already on this site, and our methodology page explains how the detector works and what its verdicts mean. If you want to see what it says about your own writing, the scanner is free to try.

Frequently asked questions

Do AI detectors falsely flag British English writing?

Not in our data. We split our 324-passage human calibration corpus by dialect and re-scored it with our shipped detector: across 85 identifiably British passages, 40 pieces of Guardian journalism and 45 passages of British classic literature, zero were falsely flagged. At that sample size the true rate is below 4.3 per cent with 95 per cent confidence. American classics also came out at zero, so the detector does not favour one dialect over the other.

What kind of writing do AI detectors falsely flag most?

In our corpus, encyclopaedic writing. Thirteen of our sixteen false positives, 81 per cent, were Wikipedia articles, a 16.2 per cent false positive rate on that source against zero on journalism, classic fiction and academic abstracts. Neutral tone, uniform sentences and absence of personal voice are Wikipedia house style, and they are also the statistical signatures detectors read as AI. A student trained to write carefully neutral formal essays inherits the same risk.

Does IsItAI work well on British English?

Yes, with one honest caveat. It falsely flagged none of the 85 British passages in our corpus, but it returned an "uncertain" verdict on 62 per cent of the Guardian journalism it scored, declining to commit rather than risking a false accusation. Abstaining is deliberate design, and we publish the abstention rate because a limitation you cannot see is worse than one you can.

Was this study independent?

No, and we say so plainly: it is our corpus and our detector, published with sample sizes, confidence intervals and the numbers that do not flatter us, including a false positive on Beowulf. We have not tested other vendors on this corpus, so it shows what our tool does with British English and what any vendor could publish if they chose to.

Try Is It AI?

Detect AI-generated content instantly. 3 free scans per day.

Scan Content Now

Free AI text check

Free, no signup

Try Now