Centre for AI-Powered Learning · White paper

AI That Teaches: When Generative AI Builds Learning That Lasts

What randomized trials show about tutors, shortcuts and defaults, and a five-rung ladder for campus-wide AI

White Paper No. 1 · 18 pages33 references3 chartsFree PDF
Get the PDF
Cover of the white paper: AI That Teaches: When Generative AI Builds Learning That Lasts

Generative AI is being licensed for whole universities and school systems, yet randomized trials show it can either build learning or replace it. This white paper reviews the 2023 to 2026 evidence on AI tutors and study aids, identifies five design conditions that separate lasting learning from assisted performance, and sets out a five-rung ladder for institutions deploying AI campus-wide.

17%Lower closed-book exam scores after plain GPT-4 practice
0.31 SDTest-score gain from teacher-guided AI tutoring in Nigeria
94%UK undergraduates using generative AI for assessed work
15%Maryland students offered a course AI tutor who used it

Key findings

  1. In a randomized trial with nearly 1,000 Turkish high school students, practicing with a plain GPT-4 chatbot raised practice scores by 48 percent but lowered closed-book exam scores by 17 percent, while a hint-giving version built on teacher solutions showed no harm.
  2. A structured AI tutor working from instructor-written solutions produced gains of 0.73 to 1.3 standard deviations over an active-learning class for 194 Harvard physics students, across two lessons.
  3. A six-week, teacher-guided AI tutoring program in Nigeria raised scores on its end-of-program test by 0.31 standard deviations, and by 0.21 on a later curricular exam, at about $48 per pupil.
  4. At the University of Maryland, a course AI tutor that most instructors left in its default direct-instruction mode lowered final grades by 0.37 standard deviations among sections of the same course, with no significant grade effect in the full sample, and only about 15 percent of students offered it used it.

The paper has 2 more findings, the full evidence review and the framework.

A chart from the paper

UK undergraduates: AI use for assessed work vs tools provided

Percent

Used generative AI for assessed workSay their institution provides AI tools0%25%50%75%100%20242025202653%89%94%9%23%38%
HEPI surveys of full-time UK undergraduates, by report year (2026 wave polled December 2025). Use is our calculation: 100% minus 'none of the above' in Figure 4 of the 2026 report (47%, 11%, 6%); HEPI's 2025 report gave 53% and 88%. Provision is 'yes' in Figure 25. Source: Higher Education Policy Institute (2026).

What is inside

  1. Executive summary and key findings
  2. Why this matters now
  3. What the trials show
  4. Where the evidence is weak or contested
  5. A practical framework: the Access-to-Evidence Ladder
  6. Case snapshots
  7. What to measure
  8. Risks and open questions
  9. Recommendations for universities, educators, edtech leaders and policymakers
  10. Methodology, limits and 33 references

One recommendation for each audience

For university leaders

Before renewing any campus-wide license, make tutoring behavior the default inside courses: hints first, grounded in instructor materials, with direct answers available only by a deliberate instructor choice.

For educators

Give the tool your own worked solutions, hints and common mistakes. Do not rely on the model to generate the answers it tutors from.

For edtech leaders

Ship tutor mode as the default and let instructors, not students, decide when direct answers are available.

For policymakers

Fund independent, multi-site trials with delayed unaided tests, including trials of the learning modes that vendors launched after the published studies.

This is a desk-research synthesis completed on 5 October 2026. We prioritized randomized trials of generative AI as a tutor or study aid in upper secondary and higher education published between 2023 and 2026, then added large surveys, official reports and institutional announcements. Every source was opened and read for this paper or reused from a fact-checked WAC Spotlight article. Company-authored studies and self-reported figures are labeled as such. WAC ran no survey, trial or interviews for this paper. The framework is our own interpretation of the evidence, not a validated instrument. The cover picture is an AI-generated illustration. Send a correction

More white papers

From the same centre

See what the Centre for AI-Powered Learning works on, with its analysis and the news in its field.

Visit the centre