WAC Signal, Episode 1: What Can a Grade Still Prove? World Assessment Council · 8 October 2026 · https://worldassessmentcouncil.org/signal/what-can-a-grade-still-prove/ THE QUESTION If a machine can write the essay, and almost everyone gets an A, what exactly is a grade still telling us? This is WAC Signal, from the World Assessment Council. Evidence and ideas for the changing world of learning. Episode one. What can a grade still prove? THE SIGNAL In the space of four academic years, three things that universities rely on to make their credentials believable came under strain at the same time. The first is authorship. In August 2026, Nanyang Technological University in Singapore confirmed it will stop using automated AI detection tools from 2027. The Singapore University of Social Sciences switched its detector off. And Australia's higher education regulator had already said, in guidance published in September 2025, that detecting generative AI use in assessment with certainty is, for now, close to impossible. The second is the grade itself. In May 2026, Harvard's Faculty of Arts and Sciences voted, four hundred and fifty-eight to two hundred and one, to limit flat A grades to twenty percent of the students in a course, plus four. The reason was simple. In 2005, A's were twenty-four percent of all grades awarded in Harvard College. By 2025, they were sixty point two percent. And the third is the return of the outside check. A growing group of selective American universities have brought back admission tests. They are still a small minority. FairTest counts one hundred and sixty-one four-year colleges requiring scores for fall 2027, against more than two thousand that remain test-optional or test-free. But the direction tells you something. These stories are usually told separately. We think they are one story. A credential is a claim. It says that a named person can do certain things, to a certain standard. That claim rests on two signals. Call the first one authenticity: the work was the student's own. Call the second one distinction: the grade separates excellent work from adequate work. Generative AI has weakened the first. Grade compression has weakened the second. And the return of testing is what institutions do when they start to doubt both. THE EVIDENCE Start with authenticity. Can software tell human work from machine work? The independent tests of the first generation of detectors are not encouraging. In one 2023 study, fourteen tools, including Turnitin, were tested on fifty-four documents. None reached eighty percent accuracy. On machine-paraphrased AI text, they got it right just fifteen percent of the time. A second 2023 study tested seven detectors on short samples, some of them lightly disguised, for example by asking the model to add a few spelling mistakes. On unaltered AI text, average accuracy was thirty-nine point five percent. After those simple tricks, it fell to twenty-two point two. The sample was small, and the authors themselves caution that the tools have changed since. The errors also fall unevenly. Researchers at Stanford ran seven detectors on essays by non-native English speakers and by American eighth-graders. The tools were near-perfect on the American essays. They wrongly flagged the second-language essays as machine-written sixty-one percent of the time. And people do no better. At the University of Reading, researchers quietly submitted AI-written exam answers under thirty-three invented student names. The markers were not told. Ninety-four percent of those scripts went unnoticed, and on average they scored higher than the real students. To be fair, detection is improving. A 2025 working paper, not yet peer reviewed, found one commercial detector with error rates close to zero on everyday writing. But that was not student coursework. And even a perfect detector answers a narrower question than a university needs answered. If AI is allowed for some purposes, a score that shows AI was involved cannot show that it was misused. Meanwhile, use is close to universal. In the UK's annual student survey, the share of undergraduates using generative AI for assessed work went from fifty-three percent, to eighty-nine, to ninety-four percent in the latest edition. The share pasting AI-generated text straight into their work rose from three percent to twelve. Most of that use is permitted, and much of it is sensible. But here is the certification problem. Work produced with a tool is weak evidence of what a person can do without it. In a randomized trial with nearly a thousand math students at one school in Turkey, open access to GPT-4 raised scores on practice problems by forty-eight percent. On the exam that followed, without the tool, those same students scored seventeen percent lower than classmates who never had it. A tutor version with safeguards largely removed the harm. Now the second signal. Distinction. Harvard is the headline, but it is not alone. In England, first-class degrees rose from under sixteen percent of awards in 2010 to a peak of nearly thirty-eight percent in 2021, before falling back to just under twenty-nine. The regulator found that about eleven points of that could not be explained by students' entry grades or subject mix. In American high schools, average grades rose over a decade while average ACT scores fell. The people who read transcripts have noticed. At Harvard, admissions deans at law and medical schools said they rely on entrance test scores more than they would like. And at several Ivy-Plus colleges, students with a top SAT score earned first-year grades almost half a point higher than similar students with a mid-range score, while a perfect high-school average predicted barely more than a good one. The pattern is the same everywhere. When an internal signal saturates, the people who rely on it go looking for an external one. THE TENSION So should every university bring back the detector, the cap and the test? The evidence says no, and it says so clearly. Detector studies age quickly, and even a small error rate, applied to tens of thousands of scripts, produces wrongly accused students. In 2025, an Australian university that had leaned on a detector stopped using it after finding it ineffective. One nursing student waited six months to be cleared. The case for tests comes mostly from highly selective colleges, where applicants' grades are bunched at the top. Research in Chicago found high-school grades predicted college graduation strongly and consistently. The evidence does not show that every institution needs a test. And caps have failed before. Princeton dropped its numerical grade targets after a decade, because they were read as quotas. At Wellesley, a cap eased grade compression, but enrollments fell in the departments it constrained, student ratings dropped, and racial gaps in grades widened. Policing, in other words, has a poor record on both signals. THE IMPLICATION If policing cannot rebuild trust, design has to. Australia's regulator has argued that assessment should be secured at meaningful points across a whole program, not in every single task. The University of Sydney has done exactly that, at a scale of more than two million submissions a year. Every task now sits in one of two lanes. Secure tasks are supervised in person. Open tasks assume AI use, and build it in. Singapore Management University puts the goal well. The most reliable assurance, its associate provost told CNA, is to design assessments that require students to show real competence, quote, not to detect whether AI was used after the fact. And Harvard has paired its cap on A grades with a percentile rank for honours, so that distinction is visible again. As the dean of undergraduate education put it, we owe our students a functioning grading system. THE WAC VIEW Rebuilding trust in a grade is mainly a design problem, and only secondarily a policing problem. In our white paper, we set out a Certification Map that program teams can apply directly. List the six to ten things a graduate is certified to do. Sort every task into a secure lane or an open lane. Certify each claim at least twice under secure conditions, with the last check close to graduation. Publish how top grades are earned, and the share of students who earn them. Three rules follow. An open task cannot carry a no-AI instruction, because nothing enforces it. A top grade should depend on at least one secure task, or it certifies the tools as much as the student. And secure does not have to mean a written exam. An oral defense, or an observed practical, is secure too, and each one certifies things an exam cannot. One honest caveat. The evidence that these redesigns improve learning is still thin. Oral and supervised formats bring costs, and fairness risks of their own. They need to be checked for second-language and disabled students, not assumed to be fair. WHAT TO WATCH One. Whether more universities follow Singapore and Australia in switching off automated detection, and what replaces it. Two. Harvard's first grade distribution under the new rules, from fall 2027, and whether students accept it where Princeton's did not. Three. The first real outcome studies of two-lane assessment. Not policy documents. Evidence on learning, on misconduct, and on fairness. The sources for every figure in this episode are listed on the episode page, along with the full white paper, What Can a Grade Still Prove, at worldassessmentcouncil.org. WAC Signal is produced by the World Assessment Council, with AI-assisted research and AI voices, and every claim is checked against its source before publication. This has been WAC Signal. Evidence and ideas for the changing world of learning. Sources: https://worldassessmentcouncil.org/signal/what-can-a-grade-still-prove/#sources WAC Signal is an editorial audio series from the World Assessment Council. Episodes are researched and written with AI-assisted editorial tools, narrated with AI voices, and checked against their cited sources before publication. Every episode page lists its sources.