What Can a Grade Still Prove?

Episode 1 · The Assessment Question
What Can a Grade Still Prove?
What Can a Grade Still Prove?
If a machine can write the essay, and almost everyone gets an A, what is a grade still telling us? Detection, grade compression and the return of admission tests are usually reported as three stories. This episode treats them as one, weighs what the evidence shows and where it is contested, and sets out what universities can do through assessment design.
- Why AI detectors and human markers both struggle to tell machine work from student work
- How grade compression at Harvard and in England has weakened what a top grade signals
- Why policing, caps and tests have a mixed record, and what has gone wrong before
- The Certification Map: secure and open lanes, and three rules for program teams
- Three things to watch next
Full transcript
The full text of the episode, so you can read it instead of, or alongside, listening. Download it as a text file.
Show or hide the transcript
The question
If a machine can write the essay, and almost everyone gets an A, what exactly is a grade still telling us?
This is WAC Signal, from the World Assessment Council. Evidence and ideas for the changing world of learning.
Episode one. What can a grade still prove?
The signal
In the space of four academic years, three things that universities rely on to make their credentials believable came under strain at the same time. The first is authorship. In August 2026, Nanyang Technological University in Singapore confirmed it will stop using automated AI detection tools from 2027. The Singapore University of Social Sciences switched its detector off. And Australia's higher education regulator had already said, in guidance published in September 2025, that detecting generative AI use in assessment with certainty is, for now, close to impossible.
The second is the grade itself. In May 2026, Harvard's Faculty of Arts and Sciences voted, four hundred and fifty-eight to two hundred and one, to limit flat A grades to twenty percent of the students in a course, plus four. The reason was simple. In 2005, A's were twenty-four percent of all grades awarded in Harvard College. By 2025, they were sixty point two percent.
And the third is the return of the outside check. A growing group of selective American universities have brought back admission tests. They are still a small minority. FairTest counts one hundred and sixty-one four-year colleges requiring scores for fall 2027, against more than two thousand that remain test-optional or test-free. But the direction tells you something.
These stories are usually told separately. We think they are one story.
A credential is a claim. It says that a named person can do certain things, to a certain standard. That claim rests on two signals. Call the first one authenticity: the work was the student's own. Call the second one distinction: the grade separates excellent work from adequate work. Generative AI has weakened the first. Grade compression has weakened the second. And the return of testing is what institutions do when they start to doubt both.
The evidence
Start with authenticity. Can software tell human work from machine work? The independent tests of the first generation of detectors are not encouraging. In one 2023 study, fourteen tools, including Turnitin, were tested on fifty-four documents. None reached eighty percent accuracy. On machine-paraphrased AI text, they got it right just fifteen percent of the time.
A second 2023 study tested seven detectors on short samples, some of them lightly disguised, for example by asking the model to add a few spelling mistakes. On unaltered AI text, average accuracy was thirty-nine point five percent. After those simple tricks, it fell to twenty-two point two. The sample was small, and the authors themselves caution that the tools have changed since.
The errors also fall unevenly. Researchers at Stanford ran seven detectors on essays by non-native English speakers and by American eighth-graders. The tools were near-perfect on the American essays. They wrongly flagged the second-language essays as machine-written sixty-one percent of the time.
And people do no better. At the University of Reading, researchers quietly submitted AI-written exam answers under thirty-three invented student names. The markers were not told. Ninety-four percent of those scripts went unnoticed, and on average they scored higher than the real students.
To be fair, detection is improving. A 2025 working paper, not yet peer reviewed, found one commercial detector with error rates close to zero on everyday writing. But that was not student coursework. And even a perfect detector answers a narrower question than a university needs answered. If AI is allowed for some purposes, a score that shows AI was involved cannot show that it was misused.
Meanwhile, use is close to universal. In the UK's annual student survey, the share of undergraduates using generative AI for assessed work went from fifty-three percent, to eighty-nine, to ninety-four percent in the latest edition. The share pasting AI-generated text straight into their work rose from three percent to twelve.
Most of that use is permitted, and much of it is sensible. But here is the certification problem. Work produced with a tool is weak evidence of what a person can do without it. In a randomized trial with nearly a thousand math students at one school in Turkey, open access to GPT-4 raised scores on practice problems by forty-eight percent. On the exam that followed, without the tool, those same students scored seventeen percent lower than classmates who never had it. A tutor version with safeguards largely removed the harm.
Now the second signal. Distinction. Harvard is the headline, but it is not alone. In England, first-class degrees rose from under sixteen percent of awards in 2010 to a peak of nearly thirty-eight percent in 2021, before falling back to just under twenty-nine. The regulator found that about eleven points of that could not be explained by students' entry grades or subject mix. In American high schools, average grades rose over a decade while average ACT scores fell.
The people who read transcripts have noticed. At Harvard, admissions deans at law and medical schools said they rely on entrance test scores more than they would like. And at several Ivy-Plus colleges, students with a top SAT score earned first-year grades almost half a point higher than similar students with a mid-range score, while a perfect high-school average predicted barely more than a good one.
The pattern is the same everywhere. When an internal signal saturates, the people who rely on it go looking for an external one.
The tension
So should every university bring back the detector, the cap and the test? The evidence says no, and it says so clearly.
Detector studies age quickly, and even a small error rate, applied to tens of thousands of scripts, produces wrongly accused students. In 2025, an Australian university that had leaned on a detector stopped using it after finding it ineffective. One nursing student waited six months to be cleared.
The case for tests comes mostly from highly selective colleges, where applicants' grades are bunched at the top. Research in Chicago found high-school grades predicted college graduation strongly and consistently. The evidence does not show that every institution needs a test.
And caps have failed before. Princeton dropped its numerical grade targets after a decade, because they were read as quotas. At Wellesley, a cap eased grade compression, but enrollments fell in the departments it constrained, student ratings dropped, and racial gaps in grades widened.
Policing, in other words, has a poor record on both signals.
The implication
If policing cannot rebuild trust, design has to. Australia's regulator has argued that assessment should be secured at meaningful points across a whole program, not in every single task. The University of Sydney has done exactly that, at a scale of more than two million submissions a year. Every task now sits in one of two lanes. Secure tasks are supervised in person. Open tasks assume AI use, and build it in.
Singapore Management University puts the goal well. The most reliable assurance, its associate provost told CNA, is to design assessments that require students to show real competence, quote, not to detect whether AI was used after the fact.
And Harvard has paired its cap on A grades with a percentile rank for honours, so that distinction is visible again. As the dean of undergraduate education put it, we owe our students a functioning grading system.
The WAC view
Rebuilding trust in a grade is mainly a design problem, and only secondarily a policing problem. In our white paper, we set out a Certification Map that program teams can apply directly. List the six to ten things a graduate is certified to do. Sort every task into a secure lane or an open lane. Certify each claim at least twice under secure conditions, with the last check close to graduation. Publish how top grades are earned, and the share of students who earn them.
Three rules follow. An open task cannot carry a no-AI instruction, because nothing enforces it. A top grade should depend on at least one secure task, or it certifies the tools as much as the student. And secure does not have to mean a written exam. An oral defense, or an observed practical, is secure too, and each one certifies things an exam cannot.
One honest caveat. The evidence that these redesigns improve learning is still thin. Oral and supervised formats bring costs, and fairness risks of their own. They need to be checked for second-language and disabled students, not assumed to be fair.
What to watch
One. Whether more universities follow Singapore and Australia in switching off automated detection, and what replaces it.
Two. Harvard's first grade distribution under the new rules, from fall 2027, and whether students accept it where Princeton's did not.
Three. The first real outcome studies of two-lane assessment. Not policy documents. Evidence on learning, on misconduct, and on fairness.
The sources for every figure in this episode are listed on the episode page, along with the full white paper, What Can a Grade Still Prove, at worldassessmentcouncil.org.
WAC Signal is produced by the World Assessment Council, with AI-assisted research and AI voices, and every claim is checked against its source before publication.
This has been WAC Signal. Evidence and ideas for the changing world of learning.
Sources for this episode
Every figure in the episode comes from these sources, as used and checked in the white paper.
- Allensworth, E. M., and Clark, K. High School GPAs and ACT Scores as Predictors of College Completion: Examining Assumptions about Consistency across High Schools. Educational Researcher, 49(3), 198-211 (abstract only, read on ERIC), 2020. Source
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O., and Mariman, R. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26) (full text read on Europe PMC), 2025. Source
- Bergin, J. University wrongly accuses students of using artificial intelligence to cheat. Australian Broadcasting Corporation, 2025. Source
- Bridgeman, A., and Liu, D. The Sydney Assessment Framework. Teaching@Sydney, The University of Sydney, 2025. Source
- Butcher, K. F., McEwan, P. J., and Weerapana, A. The Effects of an Anti-Grade-Inflation Policy at Wellesley College. Journal of Economic Perspectives, 28(3), 2014. Source
- CNA (Channel NewsAsia) SUSS drops AI detector as more Singapore universities question reliability of such tools. CNA, 2026. Source
- Claybaugh, A. (Harvard College Office of Undergraduate Education) Re-Centering Academics at Harvard College: Update on Grading and Workload. Harvard College Office of Undergraduate Education, 2025. Source
- FairTest (National Center for Fair & Open Testing) Overwhelming Majority of U.S. Colleges and Universities Remain ACT/SAT-Optional or Free for Fall 2027. FairTest, 2026. Source
- Freeman, J. Student Generative AI Survey 2025 (HEPI Policy Note 61). Higher Education Policy Institute and Kortext, 2025. Source
- Friedman, J., Sacerdote, B., and Tine, M. Standardized Test Scores and Academic Performance at Ivy-Plus Colleges. Opportunity Insights, 2024. Source
- Harvard College Office of Undergraduate Education, Subcommittee on Grading A Proposal for Updating Grading Policies. Harvard Faculty of Arts and Sciences, 2026. Source
- Harvard Faculty of Arts and Sciences Faculty decisively approve grading changes. FAS Current, 2026. Source
- Jabarian, B., and Imas, A. Artificial Writing and Automated Detection (Working Paper 34223). National Bureau of Economic Research (full text read as Becker Friedman Institute Working Paper 2025-116), 2025. Source
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., and Zou, J. GPT detectors are biased against non-native English writers. Patterns, 4(7) (read as arXiv:2304.02819), 2023. Source
- Lodge, J. M., Bearman, M., Dawson, P., Gniel, H., Harper, R., Liu, D., McLean, J., and Ucnik, L. Enacting assessment reform in a time of artificial intelligence. Tertiary Education Quality and Standards Agency (TEQSA), Australian Government, 2025. Source
- Lodge, J. M., Howard, S., Bearman, M., Dawson, P., and associates Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency (TEQSA), Australian Government, 2023. Source
- Office for Students Analysis of degree classifications over time: Changes in graduate attainment from 2010-11 to 2023-24 (OfS 2026.02). Office for Students (official statistics, England), 2026. Source
- Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., and Khuat, H. Q. GenAI Detection Tools, Adversarial Techniques and Implications for Inclusivity in Higher Education. arXiv:2403.19148 (preprint; published in revised form as 'Simple techniques to bypass GenAI text detectors: implications for inclusive education', International Journal of Educational Technology in Higher Education, 21, article 53), 2024. Source
- Princeton University, Ad Hoc Committee to Review Policies Regarding Assessment and Grading Report from the Ad Hoc Committee to Review Policies Regarding Assessment and Grading. Princeton University, 2014. Source
- Sanchez, E., and Moore, R. Grade Inflation Continues to Grow in the Past Decade (Research Report R2134). ACT Research, 2022. Source
- Scarfe, P., Watcham, K., Clarke, A., and Roesch, E. A real-world test of artificial intelligence infiltration of a university examinations system: A 'Turing Test' case study. PLOS ONE, 2024. Source
- Stephenson, R., and Armstrong, C. Student Generative Artificial Intelligence Survey 2026 (HEPI Report 199). Higher Education Policy Institute, 2026. Source
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltynek, T., Guerrero-Dib, J., Popoola, O., Sigut, P., and Waddington, L. Testing of Detection Tools for AI-Generated Text. International Journal for Educational Integrity, 19, article 26 (read as arXiv:2306.15666), 2023. Source
WAC Signal is an editorial audio series from the World Assessment Council. Episodes are researched and written with AI-assisted editorial tools, narrated with AI voices, and checked against their cited sources before publication. Every episode page lists its sources. Voices: ElevenLabs AI voices.