When AI Can Teach, What Is the Teacher For?

Episode 2 · AI & Learning
When AI Can Teach, What Is the Teacher For?
When AI Can Teach, What Is the Teacher For?
Randomized trials show the same generative AI can either build learning or replace it, and the difference depends largely on how the tool is designed and set up. This episode looks at trials at Harvard, in Nigeria, in UK schools, at the University of Maryland and in Turkey to ask what teachers contribute when AI can tutor, and sets out WAC's Access-to-Evidence Ladder for campus-wide AI.
- Why campus-wide AI licenses do not guarantee learning that lasts
- Harvard, Nigeria and UK trials where teacher-shaped AI tutoring raised later scores
- The Maryland trial where a default-mode course tutor went largely unused
- Small samples, short trials, working papers and a retracted meta-analysis
- WAC's five design conditions and the Access-to-Evidence Ladder
Full transcript
The full text of the episode, so you can read it instead of, or alongside, listening. Download it as a text file.
Show or hide the transcript
The question
If an AI tutor can take a student through a physics lesson in less time than a class takes, and the student scores higher afterward, what exactly is the teacher for?
This is WAC Signal, from the World Assessment Council. Evidence and ideas for the changing world of learning.
Episode two. When AI can teach, what is the teacher for?
The signal
Universities and school systems are now buying generative AI for everyone. In January 2025, California State University signed an eighteen-month, seventeen million dollar contract with OpenAI to give all twenty-two of its campuses ChatGPT Edu. In March 2026, Google announced that all twenty of Malaysia's public universities had enabled Gemini for Education, reaching nearly six hundred thousand students.
Students did not wait for the licenses. In the Higher Education Policy Institute's survey of full-time UK undergraduates, the share using generative AI to help with assessed work went from fifty-three percent, to eighty-nine, to ninety-four percent in the latest wave. The share saying their institution provides AI tools rose too, but only to thirty-eight percent.
The open question is whether these tools build learning that lasts once the tool is taken away. And underneath it sits another, about the people at the front of the room.
One distinction runs through this episode. Performance is what a student produces with the tool open. Learning is what the student can do afterward, alone. The OECD describes cases where the first rises and the second does not as a paradox that several studies have documented.
The evidence
Start at Harvard. In fall 2023, Greg Kestin, Kelly Miller and colleagues ran a randomized trial with one hundred and ninety-four students in an introductory physics course. Each student learned one topic from a custom AI tutor at home, and another in an active-learning class. Post-test scores were higher after the tutor, by an estimated zero point seven three to one point three standard deviations. Students got there in a median of forty-nine minutes, against sixty in class.
But look at what the tutor was. It worked from instructor-written, step-by-step solutions, and led students through each problem one part at a time. And the authors draw a clear line themselves. In their words, an AI tutor should not replace in-person teaching.
The second trial is in Nigeria. In Benin City, a World Bank team led by Martín De Simone ran a six-week after-school program of twelve ninety-minute sessions, using Microsoft Copilot, then powered by GPT-4. Each session opened with a teacher-provided prompt that positioned the model as a tutor. Students worked in pairs. Teachers circulated.
Selected students scored zero point three one standard deviations higher on the end-of-program test, which covered English, AI knowledge and digital skills, and zero point two one higher on the regular third-term English exam. And the authors credit the whole intervention, the AI plus the teacher guidance and prompts, not the chatbot alone.
A third design keeps a human between the AI and the student. In five UK secondary schools, Google's LearnLM team and the math platform Eedi ran an exploratory trial with one hundred and sixty-five students. Expert tutors supervised the model and reviewed its drafted messages, approving most with no or minimal edits. Students supported this way solved new problems on later topics sixty-six point two percent of the time, against sixty point seven percent for those tutored by humans alone.
Now the other side. The University of Maryland study is the closest available test of a campus-style rollout. In fall 2025, thirty instructors teaching two thousand three hundred and seventy-nine undergraduates were randomly assigned to have, or not have, a GPT-4o study assistant built into the learning management system.
Instructors could switch it to a tutor mode that guides students toward an answer. Most kept the default, direct-instruction mode. Only two ended the term in tutor mode. And none of the ten instructors with the tool who answered a survey had built it into their assignments.
Among sections of the same course, final grades were zero point three seven standard deviations lower with the tool, and participation in instructor-designed online activities fell by zero point nine standard deviations. Across the full sample, the authors found no significant grade effect, but a similar fall in participation.
The Turkish trial we mentioned in episode one fits the same pattern. Same GPT-4 model, two designs. The plain chatbot raised practice scores by forty-eight percent, then lowered closed-book exam scores by seventeen percent. A version that gave hints, built on teacher-written solutions and common mistakes, showed no significant difference from students who had no AI. And students using the plain version did not perceive that they had learned less.
The tension
So is the teacher the secret ingredient? The evidence cannot say that cleanly.
The positive trials are short and narrow. The Harvard study covered two lessons of introductory material. The Turkish trial ran four ninety-minute sessions in one school. In Nigeria, the control group received no extra instruction.
Who ran the study matters too. The LearnLM trial was written by Google's own team with Eedi, and is labeled exploratory. The Harvard team evaluated a tutor they had built themselves. And several of these studies, including Maryland, Nigeria and LearnLM, are working papers or preprints, not peer-reviewed articles.
The cautionary result has limits as well. Only about fifteen percent of Maryland students offered the tutor used it even once. Both groups already had ChatGPT and Gemini through the university. The outcome was course grades, not an independent test. And the authors say their findings concern the tool as configured, not more scaffolded tutoring.
Even the syntheses are not yet a safe guide. In April 2026, a journal retracted a meta-analysis of ChatGPT's effect on learning, citing discrepancies in its analysis. And claims about what AI does to the brain rest largely on one preprint with fifty-four participants.
Finally, the tools have moved on. The main trials used GPT-4 or GPT-4o.
The implication
Put the trials side by side and, in our reading, the teacher does not disappear. The teacher moves. In the designs that raised later scores, people had shaped what the tool did. They wrote the solutions it tutored from. They set the prompt that made it a tutor. They stood in the room, or between the machine and the student.
At Maryland, those choices were mostly left undone. The default stayed on, the tool sat outside the assignments, and participation in instructor-designed activities went down.
Ethan Mollick of the Wharton School, discussing these studies, put it this way: the default mode of AI is to do the work for you, not with you. A license delivers that default. Changing it, in our view, is teaching work.
We would add a job the tutor cannot do for itself, which is checking. In Turkey, students did not feel they had learned less, while their exam scores fell. Only a test taken without the tool could show it. Setting that test, and reading it honestly, is the teacher's job too.
The WAC view
Our white paper identifies five design conditions that separate the tools that taught from the tools that only helped. Hints before answers. Grounding in instructor material. A place in the course. A person in the loop. And an unaided measure of learning.
We turn those into the Access-to-Evidence Ladder. Five rungs, each with proof an institution should be able to show. Rung one is access: a license, measured by activated accounts and weekly use. Rung two is rules for each assessment, with training for staff and students. Rung three is tutor by default: hints first, grounded in course materials. Rung four is building it into teaching, with teachers reviewing the logs. Rung five is unaided proof: testing learning without the tool, and publishing the results.
Only the first rung can be delivered by a vendor alone. So before renewing any campus-wide license, make tutoring the default inside courses, with direct answers available only by a deliberate instructor choice. Write evaluation into the contract. And do not renew on account numbers alone.
For teachers: give the tool your own worked solutions, hints and common mistakes. Build it into specific assignments and class time. Ask students to attempt a problem before they ask for help. And include at least one unaided task in every unit.
One honest caveat. The ladder is our own interpretation of the evidence, not a validated instrument. And no trial reviewed in the paper follows students beyond a term.
What to watch
One. Estonia's AI Leap, which publishes its targets and its misses, including weekly use among activated accounts of twenty percent against a sixty percent target in May 2026. University of Tartu researchers are running a large randomized trial of its long-term impact.
Two. Whether campus-wide contracts start to require usage logs, mode data and the right to publish results, instead of counting seats.
Three. Independent, multi-site trials with delayed unaided tests, including trials of the learning modes vendors launched after the published studies. That is the evidence that would show what the teacher adds.
The sources for every figure in this episode are listed on the episode page, along with the full white paper, AI That Teaches, at worldassessmentcouncil.org.
WAC Signal is produced by the World Assessment Council, with AI-assisted research and AI voices, and every claim is checked against its source before publication.
This has been WAC Signal. Evidence and ideas for the changing world of learning.
Sources for this episode
Every figure in the episode comes from these sources, as used and checked in the white paper.
- AI Leap Foundation (TI-Hüpe) AI Leap: Human intelligence alongside artificial (programme overview). AI Leap Foundation, Estonia, 2026. Source
- AI Leap Foundation (TI-Hüpe) Measuring Effectiveness (short-term indicators and results, 2025/26 school year). AI Leap Foundation, Estonia, 2026. Source
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122(26), open-access full text PMC12232635, 2025. Source
- Corzo, A. Cal State struck a deal with OpenAI. Some students and faculty refuse to use it. CalMatters, 2026. Source
- De Simone, M., Tiberti, F., Barron Rodriguez, M., Manolio, F., Mosuro, W. and Dikoru, E. J. From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria. World Bank Policy Research Working Paper 11125 (May 2025; version updated December 2025), 2025. Source
- Freeman, J. Student Generative AI Survey 2025. Higher Education Policy Institute, 2025. Source
- Google Enabling Gemini for Education across all Malaysian public universities. Google Malaysia Blog, 2026. Source
- Kestin, G., Miller, K., Klales, A., Milbourne, T. and Ponti, G. AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports (Nature Portfolio), 2025. Source
- Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I. and Maes, P. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv preprint 2506.08872, 2025. Source
- LearnLM Team (Google) and Eedi AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms. arXiv preprint 2512.23633, 2025. Source
- Liu, J., Sweet, T., Chen, M. H., Engelberg, J., Masters, M. C., Clark, M., Persaud, A., Lancaster, A., Hollingsworth, J. K. and Rice, J. K. The Effects of Course-Integrated AI Tutoring on Student Performance and Engagement: A Randomized University Trial. EdWorkingPaper 26-1598, Annenberg Institute at Brown University (University of Maryland authors; version October 2026), 2026. Source
- Mollick, E. Against "Brain Damage". One Useful Thing, 2025. Source
- OECD OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education. OECD Publishing, 2026. Source
- Stephenson, R. and Armstrong, C. Student Generative Artificial Intelligence Survey 2026 (HEPI Report 199). Higher Education Policy Institute, 2026. Source
- Wang, J. and Fan, W. Retraction Note: The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis. Humanities and Social Sciences Communications 13, 528 (Springer Nature), 2026. Source
WAC Signal is an editorial audio series from the World Assessment Council. Episodes are researched and written with AI-assisted editorial tools, narrated with AI voices, and checked against their cited sources before publication. Every episode page lists its sources. Voices: ElevenLabs AI voices.