TeachersAcademic IntegrityAI DetectionUniversitiesPolicy

Why a University Told Its Staff to Stop Using AI Detectors

Mount Royal Universitys August 2026 advisory tells staff to avoid AI detectors, Turnitins included, on four grounds. Accuracy is only one of them.

Paul Byrne··10 min read


The short answer. In August 2026 Mount Royal University's Generative AI in Teaching and Learning Working Group recommended that the university "avoid the use of applications to detect AI-generated content including, but not limited to, Turnitin's AI writing detection feature for use in academic integrity processes", and said it is "proceeding with a recommendation to disable the AI detection tool in Turnitin". The conclusion is not the useful part. The test is. The working group asked four separate questions: does the tool get both kinds of error wrong, who does it harm, can a score meet the university's own standard of evidence, and what does using it do to the relationship between a teacher and a student. Accuracy is one of the four, and it is not the one that decided the case.

Most arguments about AI detectors stop at accuracy. A study says a tool is right 80% of the time, another says 60%, a vendor says 98%, and everyone argues about the number. Mount Royal University in Calgary has published something more useful than another number: a written account of how a university decided the question, with each step laid out. It is an advisory, not a regulation, and it is one Canadian institution's view. It is still the clearest four-part test we have seen for whether detection belongs in a misconduct process at all.

What Mount Royal decided

The advisory is dated August 2026 and comes from the university's Generative AI in Teaching and Learning Working Group. Its summary is one sentence: the group "recommends that Mount Royal University avoid the use of applications to detect AI-generated content including, but not limited to, Turnitin's AI writing detection feature for use in academic integrity processes."

Two things about its status matter before anyone quotes it. First, it is a recommendation. The document says the working group "is proceeding with a recommendation to disable the AI detection tool in Turnitin, but in the interim wanted to advise faculty members of the limitations of the tool". Nothing on the page says the feature has been switched off. Second, it is explicit about its basis: "Our recommendation draws on peer-reviewed studies, the official statements of comparable universities, and the vendor's own documentation."

Then it sets out the four areas, in its own words: "the tool's empirical reliability in both directions of error, and more broadly equity and inclusion concerns, fit with our evidentiary standards for misconduct, and effects on classroom culture and the instructor-student relationship."

The four-part test

1. Does it get both kinds of error wrong?

The advisory begins with accuracy, as most articles do, and insists on counting both directions. A detector "can flag work by students who did not use AI (a false positive) and it can pass work by students who did use AI (a false negative)."

On the first, it cites Weber-Wulff and colleagues, who "tested fourteen detectors, Turnitin included, in the International Journal for Educational Integrity and found that none reached eighty percent overall accuracy." On the second, it cites Perkins and colleagues, who tested detectors against AI text lightly modified in ways "students would intuitively try". In that test, "Turnitin specifically fell from 50 percent accuracy on unmodified AI text to 7.9 percent on modified text, a 42.1 percentage point drop and the largest of any tool tested."

The working group's reading of those two results together is the sharpest line in the document: "students willing to make minor edits will likely pass through the tool undetected, while students who write carefully and formally without using AI are exposed to false flags. This pattern of errors is the opposite of what a tool meant to support academic fairness should produce."

2. Who does it harm?

The second question asks who the tool is wrong about, rather than how often. The advisory cites the Liang study: seven detectors tested on essays by non-native English writers and by US eighth-graders, where "an average of 61.3 percent of the non-native essays were flagged as AI-generated, while the native-speaker essays were flagged at near-zero rates."

The point is that a tool with a modest overall error rate can still concentrate its errors on one identifiable group. That is an equity question, and an institution can answer it without settling the accuracy argument. We cover the same study, and what it means for the students most likely to be flagged, in our guide for teachers on false positives.

3. Can a score meet the standard of evidence?

Most detector debates skip this part of the test. Mount Royal's case is strongest here, and it rests on three observations.

The vendor disclaims the use. The advisory quotes Turnitin's own documentation: the model "may not always be accurate (it may misidentify human-written, AI-generated, and AI-paraphrased text), so it should not be used as the sole basis for adverse actions against a student." It adds that "since July 2024, Turnitin has gone further and stopped reporting a numerical AI score at all when the result falls between zero and twenty percent, replacing the number with an asterisk in the report", quoting the company's reason: "there is a higher incidence of false positives when the percentage is between 0 and 19."

The score cannot be examined. The advisory quotes the University of British Columbia's stated rationale for not enabling the feature, which lists among its concerns that "Results from the feature are not available to review." A number that cannot be checked is hard to put in front of an appeal panel.

The score does not mean what it appears to mean. The advisory cites a 2026 paper by Bassett and colleagues in the Journal of Higher Education Policy and Management: "Even with a generous one percent false-positive rate and ninety percent true-positive rate, the probability that an individual flagged paper is actually AI-generated ranges from below fifty percent to above ninety-five percent depending on the unknown proportion of students using AI in the cohort." In plain terms, the same score is strong evidence in a class where most students used AI and weak evidence in a class where few did, and no one marking the essay knows which class they are in.

4. What does it do to the classroom?

The fourth question is about the relationship the tool sits inside. The advisory quotes Turnitin's own blog on false positives: "if you don't acknowledge that a false positive may occur, it will lead to a far more defensive and confrontational interaction that could ultimately damage relationships with students." The working group's reading: "detection encourages a posture of suspicion, and students who know their work is being assessed by such a tool respond defensively."

It also raises a subtler point about the teacher's own judgement. Citing Bassett again, it notes that the surface features instructors are told to treat as AI hallmarks, "formulaic prose, predictable structures, lists of points, formal vocabulary", appear in AI text because they occur in the human writing the models were trained on. A teacher who has learned to see those features as suspicious is learning to be suspicious of careful writing.

What it recommends instead

The advisory does not leave the space empty. It says the university's "academic integrity response should focus on three things that the literature and our peer institutions converge on: assessment design that surfaces process and authorship (in-class writing, scaffolded drafts, oral defences); transparent statements in course outlines about what AI use is and is not permitted in each course, as our guidelines support and advocate; and academic integrity processes that emphasise instructor judgement and conversation with students rather than detector scores."

Readers of this site will recognise all three. The first is the approach one university course tested at scale, which we covered in can assessment design replace AI detection. The second is what UBC and Manchester now require in writing, covered in is AI allowed in my assignment. The third is what a New York court found missing when it annulled a finding built on a Turnitin score, covered in what a court needed beyond a 100% AI score.

How this compares with the UK

Several British universities reached a similar position earlier and with fewer words. UCL says it does not use GenAI detectors when marking. King's College London chose not to enable the Turnitin AI percentage. Manchester's guidelines say detector output "cannot currently be used as evidence of malpractice." We quote each of them from their own pages in our piece on the Dubai Accord and detection. For schools, JCQ's guidance treats a detector result as one piece of evidence inside a teacher's judgement, which we go through in what JCQ says about AI detection in coursework.

What Mount Royal adds is the reasoning. The UK statements tell you the conclusion. This one shows the working, and the working is reusable by any institution that wants to decide the question rather than inherit an answer.

Where that leaves a detection vendor

We sell an AI detector, so the honest reading of this advisory is that three of its four questions apply to us as much as to Turnitin. Our published figures are counts, not one accuracy number, and they include the cases we miss: our current model wrongly flags 3 of 400 real student essays and 4 of 361 pieces of professional writing; it catches about 98 in 100 essays copied straight from a chatbot, 69 of 70 lightly reworded ones, about 81 in 100 with typos and contractions added, about 19 in 100 written in a teenage voice, and 26 of 70 that were half AI and half the student's own work. The Perkins finding that edits defeat detectors still shows in our own numbers on the teenage-voice and half-and-half cases. All of it is on our methodology page.

We describe a result as a screening signal and never as evidence for the same reason. Mount Royal's third question, whether a score can meet an evidentiary standard, has the same answer from us as from them: on its own, it cannot. A detector can tell a teacher which essay to read more carefully and which student to talk to. The conversation decides the case.

Who wrote this, and what we sell. Is It AI is an AI-writing detector for teachers. Every quotation above is from Mount Royal University's advisory or from the institutions' own pages, linked in each case. We have no connection to Mount Royal University or to Turnitin.

Sources

  • Mount Royal University, Generative AI in Teaching and Learning Working Group, Advisory: AI Writing Detection at Mount Royal University, August 2026. All quotations, and the studies it cites (Weber-Wulff et al. 2023, Perkins et al. 2024, Liang et al. 2023, Bassett et al. 2026), are taken from this page.



Frequently asked questions

Has Mount Royal University disabled Turnitin AI detection?

Not according to the advisory itself. The August 2026 document says the working group is proceeding with a recommendation to disable the AI detection tool in Turnitin and, in the interim, advises faculty of its limitations. It is a recommendation from a working group, not a statement that the feature has been switched off.

What are the four grounds in the Mount Royal advisory?

In the advisory's own words: "the tool's empirical reliability in both directions of error, and more broadly equity and inclusion concerns, fit with our evidentiary standards for misconduct, and effects on classroom culture and the instructor-student relationship." Accuracy is one of the four.

Why does the advisory say a detector score cannot support a misconduct finding?

Three reasons it gives: Turnitin's own documentation says the score should not be the sole basis for adverse action against a student; the result cannot be reviewed; and, citing Bassett et al. 2026, the probability that a flagged paper is actually AI-generated depends on how many students in the cohort used AI, which no one marking the work knows.

What does Mount Royal recommend instead of AI detection?

Three things: assessment design that surfaces process and authorship, such as in-class writing, scaffolded drafts and oral defences; transparent course-outline statements about what AI use is and is not permitted; and integrity processes built on instructor judgement and conversation with students rather than detector scores.

Try Is It AI?

Detect AI-generated content instantly. 3 free scans per day.

Scan Content Now

Free AI text check

Free, no signup

Try Now