What Can a Grade Still Prove? Rebuilding Trust in Assessment After AI
How AI, grade compression and the return of testing are pushing universities to redesign assessment

Generative AI and decades of grade inflation have each weakened confidence in what a grade or degree certifies, and the return of admission tests at selective US universities is one sign of that doubt. This white paper reviews the evidence on detection, grade compression and testing, weighs where it is contested, and sets out a Certification Map that universities can apply program by program to rebuild trust through assessment design.
Key findings
- Seven AI detectors averaged 39.5 percent accuracy on unaltered AI text and 22.2 percent after simple evasion techniques, in a 2023 test of 114 short samples.
- Markers at the University of Reading failed to flag 94 percent of GPT-4 exam answers submitted under 33 fictitious student identities.
- A's rose from 24 percent of Harvard College grades in 2005 to 60.2 percent in 2025, and first-class degrees in England rose from 15.8 percent in 2010-11 to 28.8 percent in 2023-24.
- At Ivy-Plus colleges, students with an SAT score of 1600 earned first-year grades 0.43 points higher than similar students with 1200, while a 4.0 high school average predicted less than 0.1 points more than a 3.2.
The paper has 2 more findings, the full evidence review and the framework.
A chart from the paper
A's as a share of all Harvard College grades, 2005 to 2025
Share of grades that were A's
What is inside
- Executive summary and key findings
- Three pressures on one signal
- What the evidence shows
- Where the evidence is weak or contested
- A practical framework: the Certification Map
- Case snapshots
- What to measure
- Risks and open questions
- Recommendations for universities, educators, edtech leaders and policymakers
- Methodology, limits and 32 references
One recommendation for each audience
For university leaders
Run the Certification Map on every program within two academic years, starting where open tasks carry most of the final award grade.
For educators
Label every task as secure or open, and delete any rule about AI that you have no means of enforcing.
For edtech leaders
Publish independent, reproducible error rates, including by students' language background, and say plainly what a score cannot show.
For policymakers
Ask providers to show where each program outcome is assured under secure conditions, building on the institutional action plans that all Australian providers submitted in July 2024, without prescribing exams.
This paper is desk research. We started from earlier fact-checked articles on detection, grading and testing and widened the search to peer-reviewed studies, working papers, regulator and institutional reports, and news reporting for events, published between 2014 and September 2026. We preferred primary sources and randomized or blind designs where they exist, and included the strongest counter-evidence we found. Every source was opened; two journal articles were read as published abstracts only. Figures we calculated are marked as such. The paper reports no original survey, trial, interviews or dataset, and its framework is an untested synthesis. The cover picture is an AI-generated illustration. Send a correction
More white papers
The New Map of International Education After the Big Four Tightened
The Proof-of-Skill Gap: Why Skills-First Hiring Stalled
AI That Teaches: When Generative AI Builds Learning That Lasts
From the same centre
See what the Centre for Assessment Excellence works on, with its analysis and the news in its field.


