The number that should worry anyone buying AI grading tools
A University of Cambridge study published May 22 gives the clearest evidence yet that AI grading tools are not ready for the job many vendors are already selling them for. Researchers tested three frontier models, Claude Opus 4.6, GPT-5.4, and Gemini 3 Flash, against 761 authentic undergraduate psychology essays written by 125 students across Cambridge, Nottingham, and Manchester Metropolitan universities. Accuracy at matching the grade a human examiner assigned ranged from 63 percent at Cambridge down to 53 percent at Nottingham and just 35 percent at Manchester Metropolitan.
For anyone evaluating an AI-assisted grading product for procurement, those numbers should stop the conversation rather than season it with caveats. A coin flip does better than 35 percent, and even the best-performing university in the study saw more than a third of essays graded incorrectly by AI. This is evidence that AI grading is unreliable at the core task it is sold to perform, and it comes from a controlled academic study rather than a vendor's own benchmark, which is exactly the kind of independent validation procurement teams should be demanding before signing anything.
The models fail predictably, not randomly
The more useful finding for a buyer is the pattern behind the errors rather than the raw accuracy number. Researchers found the models were consistently biased toward middle-range grades, systematically undervaluing the strongest submissions and overvaluing the weakest ones. Dr. Deborah Talmi, the Cambridge psychologist who led the study, warned that relying on this kind of grading would produce results that are homogenised and that the approach underestimates brilliance, a bias that would compress the entire grading distribution toward the middle of the curve over enough submissions to matter at institutional scale.
Dr. Alexandru Marcoci, of the Cambridge Institute for Technology and Humanity, described the mechanism plainly: the models tend to assign middling marks to all submissions. The study also found AI models were oversensitive to linguistic features like vocabulary and sentence complexity, rewarding polished style over substantive quality, exactly the failure mode a grading tool cannot afford. A biased error is harder to catch than a random one, because it stays invisible until someone audits a large enough sample to notice the pattern, and most institutions deploying a grading tool at scale are unlikely to run an audit as rigorous as this 761-essay study before rolling it out to every course on campus and every faculty member who relies on it to save time.
Confident, verbose, and wrong
One more detail from the study should concern any buyer evaluating vendor demos: AI-generated feedback ran three to eight times longer than the feedback human graders wrote for the same essays. Length and detail read as rigor in a sales demonstration. The Cambridge data shows that length correlated with accuracy not at all. A tool that produces longer, more confident-sounding feedback while getting the underlying grade wrong more than a third of the time is a worse product than a shorter, honest one.
This is the pattern enterprise buyers should recognize from other AI procurement decisions this year: fluency reads as competence in a demo, and accuracy is a separate, harder thing to verify. The Cambridge study is one of the more rigorous validations we have seen against real graded work, spanning 87 assignments across 50 modules, and it found the fluency-accuracy gap holding across three separate frontier models, not one vendor's implementation.
Faculty are not waiting for grading tools to improve
While grading automation lags, the more consequential shift is happening on the assessment design side. A Nature report published September 1 found 94 percent of UK undergraduates now use generative AI on assessed coursework, based on a 2026 survey of 1,054 students, with 12 percent admitting to inserting AI-generated text directly into their submissions. A separate survey of more than 95,000 students across 20 US universities found 9 percent used AI on assignments despite knowing it violated the rules.
Faculty response is shifting from detection toward redesign. Daniel Silver at the University of Toronto now assigns AI agent based virtual marketplace tasks that are hard to fake with a single generated essay. Etienne Roesch at the University of Reading uses screenshot analysis tasks with narrative requirements. Yanjun Shen at Chang'an University requires students to submit their AI interaction logs alongside their work, and Kristina Ruiz-Mesa at California State University has shifted weight toward in-class presentations that a model cannot produce on a student's behalf.
The honest assessment: nobody has this solved
Nicholas Mattei, a computer scientist at Tulane University, summarized the state of the field for Nature without hedging, saying nobody has a right answer right now because things are changing rapidly. That is a more useful statement than most vendor claims we hear, because it is honest about where the technology stands rather than promising a detection or grading product that will make the problem disappear, and it matches what the Cambridge accuracy numbers already show about how far the technology has to go.
MIT linguistics instructor Nikita Bezrukov offered a telling data point: in some submissions he reviewed, half the cited reference literature did not exist in real life, a hallucination pattern that a grading tool graded on style would likely miss entirely, since a confident, well-formatted citation list reads as rigor whether or not the sources are real. Ghent University's Ruben Verborgh has shifted his course's assessment focus toward how students build sustainable web applications, a skill current AI tools do not replicate well. Each of these responses costs faculty time and course redesign effort, but each is more defensible than trusting a grading model with a documented 35 to 65 percent accuracy range to make a decision that follows a student onto a transcript.
What this means for your roadmap
If your institution or company is evaluating an AI grading or assessment scoring product, ask the vendor for accuracy data validated against real human-graded work at the scale of the Cambridge study, not a marketing claim about efficiency gains. A tool that cannot demonstrate accuracy above the 53 to 65 percent range documented here has no business anywhere near a grade that affects a student's transcript or a certification that affects someone's career.
The more durable investment, and the one the leading institutions in this reporting are already making, is in assessment redesign rather than detection or grading automation. Tasks that require live interaction, iterative process documentation, or in-person defense are harder for AI to fake and harder for a flawed grading model to misjudge. That is a curriculum and workflow investment, not a procurement one, and it is the one more likely to survive whatever the next generation of models can do.



