TeachersAssessmentAcademic IntegrityAI DetectionUniversities

Can Assessment Design Replace AI Detection? What One University Course Tried

A 130-student course made students record their reasoning before writing code. What the Ashesi case shows about proving understanding, and where detection fits.

Paul Byrne··5 min read


The short answer. Not on its own, but it changes the question in a useful way. Instead of asking "was AI used?", a well-designed assessment asks "can the student demonstrate the thinking that produced this work?". A newly published university case shows what that looks like in practice: 130 computer-science students had to explain their approach aloud, submit handwritten working, and pass through a structured mock technical interview before writing any code, generating more than 7,500 audio recordings and about 1,400 handwritten note submissions that let instructors inspect reasoning, not just output. It took real design effort and instructor oversight, and it is evidence from one course, not proof the approach generalises. The most defensible position for schools combines both: assessment that captures process, and detection used as a screening signal, never as proof.

Every conversation about AI and coursework eventually reaches the same fork. One path tries to catch AI-written work after submission. The other redesigns the work so that understanding has to be demonstrated along the way. Most of what is written about the second path is advice. On 25 August, Ashesi University published something rarer: a detailed account of actually doing it, at course scale, with numbers.

What the course actually did

The study, by researchers from St. Olaf College and the University of British Columbia in collaboration with Ashesi University, was run in a Data Structures and Algorithms course with 130 second-year computer-science students. Rather than a traditional programming assignment marked on finished code, the researchers rebuilt the exercise to resemble a technical interview.

Before writing any Java code, students had to explain how they would approach the problem, describe their reasoning, develop pseudocode, and analyse the efficiency of their proposed solution. They spoke their explanations aloud into recordings and uploaded handwritten notes to support their thinking. An instructor-designed AI bot guided them through each stage, reviewed the spoken explanations and notes, gave feedback, and prompted them to the next step. Only after completing that structured reflection did they start implementing.

The scale of the evidence trail is the striking part. Across the assignment, in the university's words, "students generated more than 7,500 audio recordings and approximately 1,400 handwritten note submissions", which were anonymised and then analysed to identify patterns in students' reasoning. The assessment itself produced a record of thinking that no detector score can offer.

Why this matters for the detection debate

AI detectors, ours included, estimate whether text carries statistical patterns consistent with AI writing. They cannot establish authorship, which is why JCQ treats detector results as one piece of evidence for a teacher's judgement, never proof, and why our own published testing reports counts rather than a single accuracy figure, including 4 of 400 real student essays wrongly flagged.

Process-based assessment attacks the problem from the other end. A student who recorded their reasoning at every stage has produced positive evidence of understanding. A student who cannot explain the work in their own recording has surfaced the problem without any classifier being involved. It converts an unanswerable question, "did AI write this?", into an observable one, "can this student walk me through their own solution?", the same logic behind the oldest advice we give teachers: compare the work against what you know of the student and talk to them about it.

What it cost, and what the study does not show

The honest version of this case includes the overheads, and the researchers are open about them. The assignment took deliberate preparation: students were introduced to the format about ten days in advance with written instructions. The AI bot needed carefully designed prompts, ongoing testing, and regular refinement, and the researchers note that instructors still had to monitor the interactions because automated systems "could occasionally produce inaccurate transcriptions or responses that do not fully align with instructional goals". Some students initially worried whether the system would understand different English accents or read their handwriting, a real concern in a classroom drawing students from across Africa.

The limits matter too. This is one assignment, in one course, in one discipline where spoken reasoning and pseudocode are natural artefacts. It does not show the method generalises to an English essay, and it does not show misconduct became impossible; a determined student could rehearse an AI-generated solution well enough to narrate it. The study's own conclusion is measured: AI worked here as a complement to the instructor, not a replacement, and "meaningful learning still depends on careful course design, active instructor oversight, and opportunities for students to demonstrate their thinking".

What schools can borrow without rebuilding every course

Few departments can rebuild an assignment as a four-stage AI-guided interview this term. But the underlying design moves scale down:

  • Ask for the middle, not just the end. Require plan, draft, and revision artefacts alongside the final piece. Version history in a document does part of what Ashesi's recordings did.

  • Make explanation routine, not an accusation. A two-minute conversation about approach, run for everyone, normalises demonstrating understanding rather than making it the penalty stage of a detector flag.

  • Put process expectations in policy before the first incident. Our school AI-use policy template covers how to define acceptable use and evidence expectations up front.

  • Keep detection in its lane. A screening scan can still tell you which submissions deserve a closer look, provided the false-positive arithmetic is understood and no score is treated as a verdict.

Detection and assessment design are not rivals; they answer different questions at different points. The Ashesi case is the best-documented evidence yet that the second question, can the student show their thinking, is answerable at course scale. If your institution is weighing where detection fits in that mix, our position is published in full on the methodology page, and you can scan work at Is It AI? knowing the result is a screening signal to start a conversation, never a finding.

Frequently asked questions

Can assessment design replace AI detectors entirely?

Not on its own. Process-based assessment produces positive evidence of understanding, but the published case is one assignment in one computer-science course, and it does not make misconduct impossible. The most defensible position combines assessment that captures process with detection used as a screening signal that starts a conversation, never as proof.

What did the Ashesi study actually do?

Researchers from St. Olaf College and the University of British Columbia, working with Ashesi University, rebuilt a Data Structures and Algorithms assignment as a mock technical interview for 130 students. Before writing any code, students explained their approach aloud, developed pseudocode, analysed efficiency, and uploaded handwritten notes, guided by an instructor-designed AI bot. The assignment generated more than 7,500 audio recordings and approximately 1,400 handwritten note submissions.

What does process-based assessment cost teachers?

Real preparation and oversight. In the published case, students were introduced to the format about ten days ahead with written instructions, the AI bot needed carefully designed prompts and ongoing refinement, and instructors monitored interactions because automated systems occasionally produced inaccurate transcriptions. The researchers conclude AI complements rather than replaces the instructor.

What can a school borrow without rebuilding every course?

Require plan, draft, and revision artefacts alongside final work; make short explain-your-approach conversations routine for everyone rather than a penalty stage; write process expectations into policy before the first incident; and keep detector results as screening signals with the false-positive arithmetic understood.

Try Is It AI?

Detect AI-generated content instantly. 3 free scans per day.

Scan Content Now

Free AI text check

Free, no signup

Try Now