Centre for Assessment Excellence

AI Detectors Can Miss Machines and Accuse Humans. Assessment Has to Change

Two Singapore universities are switching off AI detectors, as others abroad have done. Studies have found such tools missing most AI writing and misjudging second-language students. What replaces them is harder to run.

In August 2026, two Singapore universities turned away from AI detection. Nanyang Technological University (NTU) told faculty it will stop using its AI detector from 2027, and the Singapore University of Social Sciences (SUSS) switched off its university-wide Turnitin AI detector on August 11[1]. Three more, the National University of Singapore, the Singapore Institute of Technology and the Singapore University of Technology and Design (SUTD), told CNA they do not use automated detection tools. Singapore Management University (SMU) said it still uses them, but never as the sole basis for a misconduct finding[1]. NTU provost Christian Wolfrum told CNA that "what was once a reasonable safeguard has become obsolete"[1].

The retreat is not confined to Singapore. Inside Higher Ed reported in August that at least a dozen universities, including Northwestern, Georgetown and New York University, have disabled Turnitin's AI detection[3]. In September, the New South Wales Education Standards Authority changed its rules so that schools cannot rely on detection software when deciding whether a student cheated, according to Mark Bassett, who directs academic quality and integrity at Charles Sturt University[16].

In tests, most AI writing got through

Mike Perkins and colleagues ran 805 tests of seven widely used detectors on 114 short text samples in September and October 2023. On unaltered AI-generated text, the tools were right 39.5% of the time on average. Once the text was reworked with prompts designed to disguise it, average accuracy fell to about 22%, and Turnitin's dropped from 50% to 7.9%[4]. The authors caution that both the AI models and the detectors have changed since[4]. OpenAI, the maker of ChatGPT, shut down its AI text classifier in July 2023 because of its low accuracy[15].

Human markers working without detection software struggled too. At the University of Reading in 2023, Peter Scarfe and colleagues created 33 fake student accounts and submitted unedited GPT-4 answers to take-home exams in five undergraduate psychology modules. Markers were not told. They flagged 6% of the scripts; 94% went undetected, and the AI answers scored about half a grade boundary higher than real students on average[5].

National figures point the same way. A Guardian investigation based on freedom-of-information requests found almost 7,000 proven cases of AI cheating at UK universities in 2023-24, or 5.1 for every 1,000 students[6]. Yet in a December 2025 survey of 1,054 UK undergraduates for the Higher Education Policy Institute (HEPI), 12% said they put AI-generated text directly into assessed work, up from 3% in 2024[7]. The years differ, and not all AI use breaks the rules, but on our reading the distance between 5 in 1,000 and 12 in 100 suggests proven cases are a small fraction of actual use.

The false alarms fall on the wrong students

Errors in the other direction are not random. Stanford researchers Weixin Liang, James Zou and colleagues ran seven detectors on 91 TOEFL essays written by non-native English speakers and on 88 essays by US eighth-graders. The detectors were near-perfect on the American essays. On the TOEFL essays, the average false-positive rate was 61.22%: 89 of the 91 were flagged as AI by at least one detector, and all seven agreed on 18[8].

The reason is mechanical. These detectors treat predictable word choice as a sign of a machine, and people writing in a second language tend to draw on a narrower vocabulary[8]. The sample is small and dates from 2023, and GPTZero and Turnitin have published research disputing the bias, Perkins and colleagues note[4].

Small error rates scale. Turnitin said in 2023 that its false-positive rate was below 1% for documents with more than 20% AI writing, and higher when less than 20% of a document is flagged[9]. Vanderbilt University did the arithmetic before disabling the tool in August 2023: at 1%, around 750 of the 75,000 papers it sent to Turnitin in 2022 could have been wrongly labeled[10].

The best case for detection

The strongest counter-evidence is recent. In a September 2025 working paper, researchers Brian Jabarian and Alex Imas tested four detectors on a large body of human and AI text. One of them, Pangram, had false-positive and false-negative rates near zero, even on text run through tools built to disguise AI writing. It was the only detector to stay under a 0.5% false-positive cap without losing accuracy[11].

On September 29, 2026, Dartmouth's interim dean of faculty authorized professors to use Pangram, and only Pangram, on student work. Faculty must tell students whether detectors may be used and talk to a student before a score affects a grade[12].

A better detector still leaves two problems. Bassett points out that even a very low false-positive rate cannot tell a teacher how likely one flagged paper is to be AI-written, because nobody knows how many papers in the pile are human[16]. And where AI is allowed for brainstorming, drafting or editing, as it can be at SUSS, a score showing AI involvement no longer answers the misconduct question[1].

What is replacing it

The University of Sydney offers a clear model. Under a policy announced in November 2024, AI is allowed by default in assessments from Semester 1 of 2025, except in exams and in-semester tests. From Semester 2, the policy moves assessment into two lanes: secure, in-person tasks that check what students can do without aids, and open tasks where any relevant tool is fair game[13]. Pro Vice-Chancellor Adam Bridgeman's reasoning is that AI can complete any take-home assessment to a high level, so a ban is neither practical nor enforceable[13].

Singapore shows what this looks like. In one SMU module, a five-minute live pitch followed by questions is worth 30% of the grade. In another, students get the essay question two weeks ahead, may research it with AI, then write the essay in class in 15 minutes without notes[2]. Other formats include staged drafts, oral defenses and tasks that grade students on finding the errors in AI-generated summaries[2]. SUTD says it embeds AI tools in its courses and requires students to declare their use[1].

The University of the Philippines said in its 2023 AI principles that decisions about AI in teaching should start from learners' educational needs[17]. Students are noticing: 65% in the HEPI survey say assessment has changed significantly because of AI[7].

“We're still relying on the unsupervised written word as evidence of learning.”

Tricia Bertram Gallant, director of the academic integrity office, UC San Diego

None of this is free. Marc Watkins of the University of Mississippi told Inside Higher Ed he cannot fathom what it would cost a single university to scale tactics such as proctored oral exams[3]. Oral exams raise fairness questions too. Nicole Brownlie of the University of Southern Queensland cites research showing they heighten anxiety and can disadvantage students speaking a second language, the same group the Stanford study found detectors misjudging. In research on Norwegian secondary schools, the same examiners asked some students fewer than ten questions and others nearly 50[14]. Her advice is to pair a short conversation with written work, not replace it[14].

The WAC view

Detection asked software to settle a question that assessment design should have settled first. The evidence supports treating any detector score as, at most, a reason to talk to a student, never as proof, and investing instead in a few secure checkpoints where students show in person what they know. Those formats cost more and carry their own fairness risks, so universities should fund them properly and track outcomes for second-language students as closely as they now scrutinize detectors.

What to do with this

For universities

Stop treating detector scores as evidence in misconduct cases, and map each program to find the few points where learning must be verified in person. Budget staff time for those secure assessments, because that is where the real cost sits.

For educators

Tell students exactly where AI is allowed, then add one supervised element to major assignments, such as a short in-class write-up or a five-minute conversation about a draft. Use consistent questions so second-language students get an equal chance.

For edtech leaders

The demand is moving from guessing who wrote a text to supporting staged drafts, in-class writing and oral assessment at scale. Publish independent error rates, including for non-native English writers, before asking institutions to trust a score.

The events behind this article, in WAC News

Sources

  1. SUSS drops AI detector as more Singapore universities question reliability of such tools. CNA (Channel NewsAsia), August 24, 2026.
  2. 'Not about preventing AI misuse': S'pore universities move from grading essays to assessing thinking. The Straits Times, August 30, 2026.
  3. AI Detectors Are Out, New Approaches Are In. Inside Higher Ed, August 5, 2026.
  4. GenAI Detection Tools, Adversarial Techniques and Implications for Inclusivity in Higher Education (Perkins, Roe, Vu, Postma, Hickerson, McGaughran, Khuat). arXiv (published in International Journal of Educational Technology in Higher Education 21, 53), March 28, 2024.
  5. A real-world test of artificial intelligence infiltration of a university examinations system: A 'Turing Test' case study (Scarfe, Watcham, Clarke, Roesch). PLOS ONE, June 26, 2024.
  6. Revealed: Thousands of UK university students caught cheating using AI. The Guardian, June 15, 2025.
  7. Student Generative Artificial Intelligence Survey 2026 (HEPI Report 199). Higher Education Policy Institute (HEPI), March 12, 2026.
  8. GPT detectors are biased against non-native English writers (Liang, Yuksekgonul, Mao, Wu, Zou). arXiv (published in Patterns, 2023), April 6, 2023.
  9. AI writing detection update from Turnitin's Chief Product Officer. Turnitin, May 23, 2023.
  10. Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector. Vanderbilt University, August 16, 2023.
  11. Artificial Writing and Automated Detection (Jabarian and Imas). National Bureau of Economic Research (Working Paper 34223), September 2025.
  12. Dean of faculty authorizes professors to use AI detector Pangram to review student work. The Dartmouth, October 1, 2026.
  13. University of Sydney's AI assessment policy: protecting integrity and empowering students. The University of Sydney, November 27, 2024.
  14. Oral exams are making a comeback to stop AI cheating. But they have their own problems (Nicole Brownlie). The Conversation, August 17, 2026.
  15. OpenAI scuttles AI-written text detector over 'low rate of accuracy'. TechCrunch, July 25, 2023.
  16. Unis and schools are moving away from AI-detection software. They should stop using it altogether (Mark A. Bassett). The Conversation, September 23, 2026.
  17. University of the Philippines Principles for Responsible and Trustworthy Artificial Intelligence. University of the Philippines, July 11, 2023.

Written by the WAC editorial team. Facts are drawn from the numbered sources above and were checked against them before publication; opinions appear only under “The WAC view”. The organisations and people named are not affiliated with the World Assessment Council unless stated. Send a correction

More from Spotlight

Assessment Excellence

Harvard Capped the A. History Says the Hard Part Comes Next

After six in ten grades became A's, Harvard's faculty voted 458 to 201 to ration the top mark from fall 2027. Princeton and Wellesley tried to hold grades down before, and both gave up.

· 6 min read
Assessment Excellence

The Test Is Back: Why Elite Universities Stopped Trusting Grades Alone

All eight Ivy League universities have now decided to require the SAT or ACT again. Dartmouth's data suggest test-optional admission hurt some disadvantaged applicants, yet FairTest counts nine in ten US colleges as still not requiring scores.

· 6 min read
AI-Powered Learning

AI Tutors Have Beaten Good Teaching and Cut Grades. Design Decides Which

Malaysia's 20 public universities have enabled Gemini for nearly 600,000 students, Google says. Randomized trials show AI tutors can double learning gains or cut exam scores, depending on how they are set up.

· 6 min read

Work with the WAC centres

Universities, educators and education companies can register their interest in joining the World Assessment Council network.

Join WAC