AI Tutors Have Beaten Good Teaching and Cut Grades. Design Decides Which
Malaysia's 20 public universities have enabled Gemini for nearly 600,000 students, Google says. Randomized trials show AI tutors can double learning gains or cut exam scores, depending on how they are set up.
On March 9, 2026, Google announced that all 20 of Malaysia's public universities had enabled Gemini for Education, reaching nearly 600,000 students and 75,000 faculty members[1].
The move takes a campus pattern to national scale. When Arizona State University announced a collaboration with OpenAI in January 2024, President Michael Crow said the university was optimistic that AI tools could help students learn faster and grasp subjects in more depth[2]. Chulalongkorn University unveiled ChulaGENIE for more than 50,000 students, faculty and staff that November, and IE University said in February 2025 that it would make ChatGPT Edu available across its community[3][4].
Students were already there. A HEPI survey of 1,054 full-time UK undergraduates, published three days after the Malaysian announcement, found that 94% use generative AI to help with assessed work, while only 38% say their institution provides AI tools[5]. The open question is no longer access but learning, and randomized trials point in opposite directions.
The trial where AI beat good teaching
At Harvard, Greg Kestin, Kelly Miller and colleagues ran a randomized trial in fall 2023 with 194 students in an introductory physics course. Each student learned one topic at home from a custom AI tutor and another in an active-learning class, a format already shown to beat passive lectures[6].
The AI lessons won. Median post-test scores were 4.5 after the tutor and 3.5 after class, from a baseline of 2.75, so median learning gains more than doubled. Students also needed less time: a median of 49 minutes, against 60 in class. The authors estimate the effect at 0.73 to 1.3 standard deviations[6].
This was not an off-the-shelf chatbot. Instructors wrote step-by-step solutions into the prompts so the tutor did not rely on GPT-4 alone to work out answers, and the platform led students through each problem one part at a time. The trial covered two lessons of introductory material, and the authors write that "an AI tutor should not replace in-person teaching"[6].
The trials where scores fell
Hamsa Bastani of the Wharton School and colleagues ran a randomized field experiment with nearly 1,000 students at one high school in Turkey. Some got GPT-4 during math practice, in one of two forms. One mimicked the standard ChatGPT interface. The other, called GPT Tutor, was told to give hints instead of answers and was supplied with teacher-written solutions and common student mistakes[7].
In practice sessions, scores were 48% higher with the plain version and 127% higher with the tutor than in the no-AI control group. Then students sat a closed-book exam. Those who had practiced with the plain version scored 17% lower than the control group. The tutor version removed that harm but produced no gain. Students in the plain-chatbot group did not perceive that they had learned less[7].
The most common message students sent the plain chatbot was a request for the answer, and when the researchers tested it on the practice problems, it answered correctly only 51% of the time[7].
Higher education now has a cautionary result of its own. A University of Maryland working paper dated October 2026 reports a randomized trial with 2,379 undergraduates and 30 instructors in fall 2025. Instructors were randomly assigned a GPT-4o tutor, built into the learning platform and drawing on course materials they selected. Among sections of the same course, final grades ended 0.37 standard deviations lower where the tutor was offered than where it was not. Across the full sample, the authors found no significant evidence of a grade effect. Recorded participation on the platform fell by about 0.9 standard deviations in both samples, and estimated grade losses were more than twice as large for first-generation students[8].
Only about 15% of students offered the tutor ever used it, both groups already had ChatGPT and Gemini through the university, and the outcome is course grades, not an independent test. The authors say the findings concern the tool as configured, not more scaffolded AI tutoring[8].
What separates the two
Across these studies, four design choices stand out.
- Who holds the answer. The Harvard tutor and GPT Tutor worked from solutions written by teachers. The plain chatbot generated its own and was wrong about half the time[6][7].
- What the default is. Maryland instructors could switch the tool to a tutoring mode that guides students toward an answer. Most kept the default direct-instruction mode, and only two ended the term in tutoring mode. The authors classed 73.8% of student interactions as direct-answer requests[8].
- What the tool displaces. At Maryland, the participation that fell included discussion responses, assignment submissions and quizzes, and none of the 10 tutor-group instructors who answered a survey reported building the tutor into assignments[8].
- Whether a person is in the loop. Tutor CoPilot, built by Stanford researchers, gives AI suggestions to human tutors, not to students. In a trial with 900 tutors and 1,800 K-12 students, those whose tutors had the tool were 4 percentage points more likely to master topics, and 9 points among students of lower-rated tutors[9].
“the default mode of AI is to do the work for you, not with you”
Discussing the Turkish and Harvard studies in 2025, Mollick argued that the way AI is used, not the fact of using it, decides whether it helps or harms learning[13].
The 2 sigma promise, checked
The case for AI tutors often leans on Benjamin Bloom's 1984 claim, as Paul von Hippel of the University of Texas at Austin recounts it in Education Next, that tutoring could raise achievement by two standard deviations, from the 50th percentile to the 98th. Khan Academy founder Sal Khan promoted the launch of the Khanmigo tutor in a 2023 talk titled The Two Sigma Solution[11].
Von Hippel traced Bloom's figure to the dissertations of two of his PhD students: three-week experiments on unfamiliar topics, scored on narrow tests. A 2020 meta-analysis he cites found an average tutoring effect of 0.37 standard deviations; none of the 96 studies it reviewed reached two[11].
The chatbot research base is unsettled too. A 2025 meta-analysis of 51 studies that reported a large positive effect of ChatGPT on learning performance was retracted on April 22, 2026. Its page still showed 447 citations in October 2026[12].
What it means for a system-wide rollout
Malaysia's package contains some of the right parts. Google says the Gemini model offered there is infused with LearnLM, its models tuned for learning, and includes Guided Learning, a feature meant to build understanding instead of handing over quick answers[1]. In an exploratory trial with 165 UK students, run by Google's LearnLM team and Eedi, students guided by LearnLM did at least as well as those with human tutors, though expert tutors supervised every message it drafted[10].
But the announcement presents Guided Learning as something students can explore, and it reports training counts and classroom examples, not test results[1]. The evidence suggests the optional path is the weak point: most of Maryland's instructors kept the default, and Turkey's students asked for the answer[7][8].
Three checks follow for any large deployment: measure what students can do unaided, make tutoring the default inside courses, and ground it in instructor-verified materials[6][7][8].
The WAC view
Switching on AI for a university system is a procurement decision. Whether it teaches is a course-design decision, and the trials suggest the second matters more. We think every large deployment should publish at least one measure of what students can do with the tool switched off. Until then, adoption numbers show how many students have an assistant, not how many have a tutor.
What to do with this
For universities
Treat a campus-wide AI license as the start of a course-design project: set tutoring behavior as the default inside courses and track at least one unassisted measure of learning, not only adoption.
For educators
Supply the tool with your own verified solutions and hints, build it into specific assignments, and keep some assessed work AI-free so you can see what students can do alone.
For edtech leaders
Ship the tutoring mode as the default, not an option, and publish results on unassisted performance; in the Maryland trial, most instructors kept the default direct-instruction mode.
The events behind this article, in WAC News
Sources
- Enabling Gemini for Education across all Malaysian public universities. Google Malaysia Blog, March 9, 2026.
- A new collaboration with OpenAI charts the future of AI in higher education. ASU News (Arizona State University), January 18, 2024.
- Chula Pioneers Responsible Use of Generative AI for Higher Education in Thailand with the Inauguration of 'ChulaGENIE,' in Collaboration with Google Cloud. Google Cloud Press Corner, November 27, 2024.
- IE University becomes one of the first top universities to integrate OpenAI tools throughout its entire academic ecosystem. IE University, February 18, 2025.
- Student Generative Artificial Intelligence Survey 2026 (HEPI Report 199), by Rose Stephenson and Charlotte Armstrong. Higher Education Policy Institute (HEPI), March 12, 2026.
- AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting (Kestin, Miller, Klales, Milbourne, Ponti). Scientific Reports (Nature Portfolio), June 3, 2025.
- Generative AI without guardrails can harm learning: Evidence from high school mathematics (Bastani, Bastani, Sungu, Ge, Kabakci, Mariman), PNAS 122(26), open-access full text PMC12232635. Proceedings of the National Academy of Sciences (PNAS), June 25, 2025.
- The Effects of Course-Integrated AI Tutoring on Student Performance and Engagement: A Randomized University Trial (Liu, Sweet, Chen and colleagues, University of Maryland), EdWorkingPaper 26-1598, version October 2026. EdWorkingPapers, Annenberg Institute at Brown University, October 2026.
- Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise (Wang, Ribeiro, Robinson, Loeb, Demszky, Stanford University). arXiv, October 3, 2024.
- AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms (LearnLM Team, Google, and Eedi). arXiv, December 29, 2025.
- Two-Sigma Tutoring: Separating Science Fiction from Science Fact, by Paul T. von Hippel. Education Next, March 7, 2024.
- RETRACTED ARTICLE: The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis (Wang and Fan; article page, marked retracted on 22 April 2026). Humanities and Social Sciences Communications (Springer Nature), May 6, 2025.
- Against "Brain Damage". One Useful Thing (Ethan Mollick), July 7, 2025.
Written by the WAC editorial team. Facts are drawn from the numbered sources above and were checked against them before publication; opinions appear only under “The WAC view”. The organisations and people named are not affiliated with the World Assessment Council unless stated. Send a correction
More from Spotlight
EdTech Funding Fell 89%. What Still Gets Funded Looks Very Different
Reach Capital's new $265 million fund is not a return to 2021. Global EdTech venture funding has fallen by almost nine-tenths, and what remains favors AI tools sold to employers and institutions that can show results.
Harvard Capped the A. History Says the Hard Part Comes Next
After six in ten grades became A's, Harvard's faculty voted 458 to 201 to ration the top mark from fall 2027. Princeton and Wellesley tried to hold grades down before, and both gave up.
Companies Dropped the Degree Requirement. Most Hired Graduates Anyway.
Coursera and Udemy have merged on the promise of a skills-first job market. Hiring records show employers changed their job ads far more than their hires, and the missing piece is trusted proof of skill.
Work with the WAC centres
Universities, educators and education companies can register their interest in joining the World Assessment Council network.