Does Turnitin Use Student Essays to Train AI? What the 2026 Policy Says
Southampton will not renew Turnitin beyond 2026-27 over proposed data terms. What integrity tools can do with student work, and what institutions should ask.
The short answer. For its Clarity AI assistant, Turnitin says the pre-trained models behind it do not use or retain student or assignment data for AI training, while anonymised assignment data is used to evaluate and improve the assistant. For its detection model, Turnitin says it is trained on a representative sample including authentic academic writing and its own in-house corpus of academic data, without stating whether that corpus includes student submissions. Both statements are live. The question for every institution is not only whether the tool detects AI but what happens to student text after submission.
When the University of Southampton announced it would not renew its Turnitin contract beyond the 2026-27 academic year, the reason, reported by Times Higher Education, was proposed terms that would have allowed student submissions to be used in AI-related service improvement. Southampton said it understood the proposed changes had been postponed, but that Turnitin still intends to revisit them, and it began phasing out immediately.
That decision is worth examining carefully, because it is not a finding about current practice. It is a signal that a third question has entered the procurement conversation around academic integrity software, alongside the two that were already there.
What are the three distinct questions about academic integrity tools?
They get conflated in staffroom conversations and sometimes in vendor marketing. Separating them makes it easier to ask suppliers the right things.
Detector performance is whether the tool correctly identifies AI-written text and correctly clears human-written text. The key failure mode to watch is false positives: a human student wrongly flagged pays a higher cost than AI-generated text that slips through. Published research by Liang et al. (Stanford, 2023) found that AI detectors flag non-native English writers at substantially higher rates than native speakers, sometimes above 60 percent, because second-language phrasing can resemble the uniform prose that detectors are trained to catch. Our own published testing held Is It AI?'s false-positive rate at 4 of 400 real student essays wrongly flagged; we publish those figures because an institution cannot make a fair judgement on a number it has not been given.
Evidentiary use is what a score can establish in a misconduct proceeding. JCQ, the UK examinations body, is explicit on this: a score is one input inside a holistic assessment, not a finding on its own. A separate post covers what JCQ actually says about AI detection in coursework. The short version: JCQ names Turnitin as one of four example programs, then states immediately that the list "is for information purposes only and does not constitute an endorsement." A score should prompt a conversation and a review of drafts, not a referral in isolation.
Data policy is what happens after submission: where student text is stored, how long it is retained, whether it can be used to improve the software, and whether the institution agreed to that use when signing the contract. This is the question Southampton raised.
What does Turnitin actually say about student data?
Turnitin's administrator FAQ for its Clarity product states that the assistant is built on pre-trained models "which do not use or retain any student or assignment data submitted to Turnitin for AI training and evaluation", and in the same answer says: "We do use anonymized student assignment data to evaluate, improve, and extend the quality of the Turnitin AI assistant responses and capabilities." Both halves are part of the same public answer, and they refer to different activities: the underlying models are not trained on student data, while anonymised student data is used to improve the assistant built on top of them.
The detection model is described separately. Turnitin's AI writing detection FAQ says the detector is trained on "a representative sample of data spread over a period of time, that includes both AI generated and authentic academic writing across geographies and subject areas", and elsewhere that its models are "trained on our vast inhouse corpus of academic data". Those statements do not say whether that corpus includes student submissions, and that unanswered question is precisely the kind of ambiguity the procurement questions below are designed to resolve.
In mid-2026, Turnitin announced licence updates that would have required institutions to consent to broader data use for AI model improvement. Following pushback from multiple institutions, Turnitin paused the rollout and said it would work with customers before making any future changes. The company has confirmed it intends to revisit the terms.
This is not the same as saying student work is currently being used to train AI models. It is saying the policy landscape is unsettled and the contract signed last year may not reflect what a renewal would require.
What questions should every institution be asking?
Whether an institution uses Turnitin or any other academic integrity service, the data-governance question now sits alongside the detection-accuracy question. The post on what happens to text you paste into an AI detector covers per-tool retention and training policies in detail, and who owns the work you submit to Turnitin covers the licence and deletion terms in the user agreement itself. For institution-level contracts, the questions that carry most weight are:
- What is stored, and for how long? Does student text enter a shared similarity repository that other institutions can match against? If so, what are the deletion terms?
- Is student text used to train, evaluate, or improve any AI system? This includes the detection model, any writing-assistant product, and any downstream service.
- Who gave consent, and when? In most deployments, students submit via an LMS integration and do not see vendor terms directly. Is the institution named as data controller, and does the data-processing agreement cover AI-related use?
- Does the current contract match what was originally signed? Turnitin's mid-2026 proposed changes show that licence terms can shift during a multi-year agreement. Checking the current version against the signed version is not excessive caution; it is basic contract management.
- What is the institution's own retention practice? Even where vendor terms are clear, an LMS gradebook may be retaining submissions indefinitely. That is an institutional decision, not a vendor one.
Why is the Southampton decision significant?
Not because it settles anything, but because of the type of institution making it. Most individual teachers lack the resource to interrogate a vendor's licence terms; many smaller schools rely on guidance from a local authority or multi-academy trust. A large research university with legal resource and a data-protection team making an explicit, public, data-governance decision raises the floor for every other buyer.
According to Times Higher Education's reporting, Southampton's reasoning was forward-looking: even though the proposed changes had been temporarily paused, the university judged that Turnitin still intends to revisit them. The decision to begin phasing out immediately, rather than waiting for the contract to expire, reflects an assessment of what renewal would likely involve rather than a finding about current practice.
Other institutions will reach different conclusions on the same facts. The Southampton decision makes the question harder to defer; it does not answer it for anyone else.
Does this apply to schools as well as universities?
Secondary schools in England and Wales are data controllers under UK GDPR for student work submitted via any platform. As our JCQ coursework guidance checklist notes, the data-governance question applies before any scan takes place: student work is personal data, and the institution is the controller.
For schools using Turnitin through a local-authority or MAT agreement, the data-processing agreement at the top of the chain governs what can happen with student submissions. If that agreement was drafted before AI-related processing was a live concern, it may need reviewing with a data-protection officer.
What about students?
If your institution uses an academic integrity platform, your submitted work sits on third-party servers under that platform's retention and data-use terms, almost certainly terms you did not see directly. That is standard for any cloud-based institutional service; the governance question sits with the institution, not the individual student.
What students do control is whether they declare AI use correctly under their institution's policy and keep the supporting materials JCQ requires: a screenshot of the AI output, the name of the tool, the date it was generated. If you are unsure what your institution allows, the post on whether it is safe to use AI for your essay covers JCQ's requirements in plain English.
If you want to check your own writing before submission, you can scan it at Is It AI?. The tool does not store the text you submit; full details are on the methodology page.
---
Sources: Times Higher Education, "Southampton dumps Turnitin over use of students' work to train AI," August 2026. Turnitin Guides, "FAQs for administrators using Turnitin Clarity" and "Turnitin's AI writing detection capabilities FAQs," both retrieved 25 August 2026. University of Sussex staff notice on the Turnitin EULA pause, August 2026. Liang et al., "GPT detectors are biased against non-native English writers," Patterns / Stanford HAI, 2023. JCQ, "AI Use in Assessments: Your role in protecting the integrity of qualifications," revision two, 30 April 2025.
Frequently asked questions
Does Turnitin currently use student essays to train AI?
For the Clarity AI assistant, Turnitin says the pre-trained models behind it do not use or retain student or assignment data for AI training, while anonymised assignment data is used to evaluate and improve the assistant. For the detection model, Turnitin describes training on authentic academic writing and an in-house corpus without stating whether student submissions are part of it. Turnitin proposed broader consent requirements in mid-2026, paused the rollout after pushback, and has confirmed it intends to revisit the policy.
Did Southampton say Turnitin is misusing student data?
No. Southampton's stated reason was concern about proposed future terms, not a finding about current data use. The university said it understood the proposed changes had been postponed but did not want to renew under terms where those changes were still planned.
Is Is It AI? storing the text I submit?
No. The tool holds text in memory for the scan, uses it to produce a screening result, and does not store it afterwards. A one-way hash with scores and passage positions is kept for up to 24 hours to avoid re-charging identical requests; the text itself is never retained. This is set out in full on the methodology page.
Should my school stop using Turnitin?
That is an institutional decision that should involve your data-protection officer and, where relevant, your MAT or local authority. Southampton's decision is a legitimate governance position taken by a large institution with the legal resource to make it. Smaller institutions should start by confirming what their current data-processing agreement actually covers.