The Return of the Oral Exam: Live Assessment at Scale
What vivas, oral defenses and AI examiners can certify, what they cost and how to run them fairly

Universities are bringing back oral exams and oral defenses because unsupervised written work no longer shows who did the thinking. This paper reviews whether live assessment is valid, reliable and fair, what it costs in examiner time, and how it has been run for hundreds of students, with an early look at voice AI. It proposes a Live Check Ladder for deciding where a live check is worth the cost.
Key findings
- In a December 2025 survey of 1,054 UK undergraduates, 94 percent used generative AI for assessed work and 65 percent said assessment had changed significantly in response to AI.
- In a randomized comparison of 99 biology students, those examined orally scored significantly higher than those who answered the same questions in writing.
- Among 896 candidates in a UK medical membership exam, two 20-minute structured orals reached a case-to-case reliability of 0.65, and only 41 percent of score variance reflected differences between candidates.
- In a three-year project with 4,020 engineering students at UC San Diego, 56.1 percent of survey respondents expected excessive stress from oral exams, but 24.1 percent later reported it, against 62.3 percent for written exams.
The paper has 2 more findings, the full evidence review and the framework.
A chart from the paper
Oral exam stress at UC San Diego: expected versus reported
Share of survey respondents
What is inside
- Executive summary and key findings
- Why live assessment is back
- What the evidence shows
- Where the evidence is weak or contested
- A practical framework: the Live Check Ladder
- Case snapshots
- What to measure
- Risks and open questions
- Recommendations for universities, educators, edtech leaders and policymakers
- Methodology, limits and 32 references
One recommendation for each audience
For university leaders
Decide program by program which claims need a live check, and use the lowest rung of the Live Check Ladder that certifies each one.
For educators
Use a question bank with a random draw, scripted hints and a published rubric that marks reasoning, not fluency.
For edtech leaders
Before selling an AI examiner, publish its agreement with blind human marking and its results by student group, including second-language speakers.
For policymakers
Recognize structured, recorded orals as a legitimate secure format in quality frameworks, and ask providers how they assure consistency and reasonable adjustments.
This paper is desk research. We began with regulator guidance and recent news reporting, then searched for peer-reviewed studies, conference papers, working papers and institutional guidance on oral assessment published between 2003 and September 2026. We preferred randomized or large-sample designs where they exist and included evidence that cuts against the format. Every source was opened; ten journal articles and working papers were read as published abstracts only. Figures we calculated are marked. The paper reports no original survey, trial, interviews or dataset, and its framework is an untested synthesis. The cover picture is an AI-generated illustration. Send a correction
More white papers
What Can a Grade Still Prove? Rebuilding Trust in Assessment After AI
The New Map of International Education After the Big Four Tightened
The Proof-of-Skill Gap: Why Skills-First Hiring Stalled
From the same centre
See what the Centre for Assessment Excellence works on, with its analysis and the news in its field.


