AI That Teaches: When Generative AI Builds Learning That Lasts
What randomized trials show about tutors, shortcuts and defaults, and a five-rung ladder for campus-wide AI

Generative AI is being licensed for whole universities and school systems, yet randomized trials show it can either build learning or replace it. This white paper reviews the 2023 to 2026 evidence on AI tutors and study aids, identifies five design conditions that separate lasting learning from assisted performance, and sets out a five-rung ladder for institutions deploying AI campus-wide.
Key findings
- In a randomized trial with nearly 1,000 Turkish high school students, practicing with a plain GPT-4 chatbot raised practice scores by 48 percent but lowered closed-book exam scores by 17 percent, while a hint-giving version built on teacher solutions showed no harm.
- A structured AI tutor working from instructor-written solutions produced gains of 0.73 to 1.3 standard deviations over an active-learning class for 194 Harvard physics students, across two lessons.
- A six-week, teacher-guided AI tutoring program in Nigeria raised scores on its end-of-program test by 0.31 standard deviations, and by 0.21 on a later curricular exam, at about $48 per pupil.
- At the University of Maryland, a course AI tutor that most instructors left in its default direct-instruction mode lowered final grades by 0.37 standard deviations among sections of the same course, with no significant grade effect in the full sample, and only about 15 percent of students offered it used it.
The paper has 2 more findings, the full evidence review and the framework.
A chart from the paper
UK undergraduates: AI use for assessed work vs tools provided
Percent
What is inside
- Executive summary and key findings
- Why this matters now
- What the trials show
- Where the evidence is weak or contested
- A practical framework: the Access-to-Evidence Ladder
- Case snapshots
- What to measure
- Risks and open questions
- Recommendations for universities, educators, edtech leaders and policymakers
- Methodology, limits and 33 references
One recommendation for each audience
For university leaders
Before renewing any campus-wide license, make tutoring behavior the default inside courses: hints first, grounded in instructor materials, with direct answers available only by a deliberate instructor choice.
For educators
Give the tool your own worked solutions, hints and common mistakes. Do not rely on the model to generate the answers it tutors from.
For edtech leaders
Ship tutor mode as the default and let instructors, not students, decide when direct answers are available.
For policymakers
Fund independent, multi-site trials with delayed unaided tests, including trials of the learning modes that vendors launched after the published studies.
This is a desk-research synthesis completed on 5 October 2026. We prioritized randomized trials of generative AI as a tutor or study aid in upper secondary and higher education published between 2023 and 2026, then added large surveys, official reports and institutional announcements. Every source was opened and read for this paper or reused from a fact-checked WAC Spotlight article. Company-authored studies and self-reported figures are labeled as such. WAC ran no survey, trial or interviews for this paper. The framework is our own interpretation of the evidence, not a validated instrument. The cover picture is an AI-generated illustration. Send a correction
More white papers
The New Map of International Education After the Big Four Tightened
The Proof-of-Skill Gap: Why Skills-First Hiring Stalled
What Can a Grade Still Prove? Rebuilding Trust in Assessment After AI
From the same centre
See what the Centre for AI-Powered Learning works on, with its analysis and the news in its field.


