Chat-based exams that grade what students can explain in their own words.
Students explain a concept to an AI interviewer in chat, the way they would in an oral exam with a TA. A second AI holds your rubric and grades the transcript. Use it as a low-stakes check after a unit, a lab sign-off, or an exam question in PrairieLearn. It works for anything a student can explain in prose: a concept, a mechanism, a design decision, a reading. It is not for math derivations or writing code.
The interviewer never sees the rubric or the answer key, so it cannot be talked into revealing them. And credit goes only to what the student says on their own: an answer the interviewer had to point at earns at most partial credit.
A simulated student taking the library assessment "The Turing Test", graded by the same models students get. Watch the second criterion: it drops to "partially met" when the interviewer has to ask, and recovers only when the student develops the idea unprompted.
How It Works
Two Agents, One Conversation
The interviewer talks to the student. After every message the evaluator re-reads the conversation, updates each criterion, and tells the interviewer what to probe next, never what the answer is. One agent that knew the answers could be argued into giving them away; this one cannot.
No Credit for Hints
A criterion counts as met only when the student volunteered it; something said after the interviewer pointed at it is partially met at best. Confidence, argument, and claims of authority do not move the grade, and student messages are treated as answers, never as instructions. Time limits are advisory and students can consult whatever they like mid-conversation; hint-fed and pasted answers read as prompted.
Authored With Your Own AI
Write the assessment by chatting with the AI you already use (claude.ai, ChatGPT, Claude Code, Cursor, or any AI tool that can connect to outside services over MCP); it connects to your account here. The server checks your draft for leaks and grading problems, then simulates good, weak, and adversarial students against it before any real student sees it. Budget about an hour for a first assessment.
Delivered in PrairieLearn
One element and one question file, no changes to PrairieLearn itself. Grades land in the gradebook; transcripts are kept for review and appeals. PrairieLearn is the only delivery surface during the beta.
Evidence
The approach ran for a semester in a University of Illinois course and is written up in a manuscript submitted to SIGCSE TS 2027. It was a small deployment, and we say so: the numbers below are what we can defend.
0 leaks in 619 automated checkseleven simulated students, six adversarial, trying to extract answers or inject instructions; none got rubric material out of the interviewer
41 real conversationseleven students, three assessments, Spring 2026; median lengths of 19–30 minutes across the three assessments, median seven student messages
On grading reliability: two AI graders re-grading the same ten transcripts blind agreed with each other exactly on 7 and within one level on 9; the two faculty authors, grading the same transcripts, agreed on 6 and 9. A ten-transcript sample proves consistency, not correctness, which is why every grade ships with its reasoning and the transcript. The evaluator deployed that semester was a weaker model that matched a blind judge exactly on only 14 of 40 attempts; the current evaluator matches the judge on every test run, which is why every model now passes an adversarial bake-off before we use it. The results are public.
Cost
About half a cent per student-minute on the current models; the table below works it out for common exam lengths. Free while we run the beta; if a price follows, it will look like those numbers rather than a site license.
Access
You do not need your own AI account; we host the models during the beta. Illinois faculty sign in with their campus account. Educators elsewhere can email for an invitation; see Get Started below.
Accuracy
Every grade is auditable: the evaluator cites which rubric level it applied for each criterion, and the full transcript sits in PrairieLearn's submission history. If you disagree with a grade, change the score there like any manually graded question.
What It Costs
Cost scales with how long students talk. From our measurements (1.4¢ of model time per student turn on the current models, and students averaging about 0.35 turns per minute in real use), a conversation costs about 0.5¢ per minute. Very long exams run slightly higher, because each reply carries more history.
Exam length
Per student
Class of 30
Class of 300
30 minutes
15¢
$4.38
$43.81
60 minutes
29¢
$8.76
$87.61
120 minutes
58¢
$17.52
$175.23
minutes
—
—
students: —
Model list prices, current as of August 2026. The paper's Spring 2026 deployment, on older models, cost a median of $0.27 per conversation.
Get Started
Illinois faculty: sign in with your campus account and request access; once approved, adding the ready-made question to a course takes about fifteen minutes, and authoring your own assessment about an hour. Everyone else: email for an invitation and say a little about the course you have in mind; expect a reply within a few days.
What leaves your LMS: the student's messages, and nothing else unless your course opts in to sending names so attempts can be matched to students. Transcripts and grades are stored here and in PrairieLearn's submission history, and are visible to you and to us, not to other instructors. Details on the PrairieLearn page.