AI Is About to Change What a Test Question Can Actually Measure
AI & ML

AI Is About to Change What a Test Question Can Actually Measure

An assessment researcher says AI can now score a Socratic back-and-forth as reliably as a bubble sheet, which turns the next assessment contract into a data governance decision as much as a testing-vendor renewal.

PublishedOctober 9, 2026
Read time5 min read
Share

The pitch is bigger than faster grading

When district technology leaders hear "AI in assessment," most still picture a faster Scantron. Temple Lovelace, executive director of the Assessment for Good program at the Advanced Education Research and Development Fund, is describing something larger. Her organization builds formative assessments for students ages 8 to 13, and she told K-12 Dive that AI now touches nearly every stage of the process: generating inventive question formats, scoring responses in something close to real time, and measuring durable skills like critical thinking that a fill-in-the-bubble test was never built to capture.

The shift from static test items to interactive ones is the part worth sitting with. "The learner talks, and you assess what they're thinking in a Socratic conversation back and forth," Lovelace said, describing formats her team has used for the past two and a half to three years. Instead of a student clicking an answer, they converse with a chatbot or upload work samples, and the system extracts evidence of reasoning and collaboration from that exchange. That is a materially different data object than a scored answer sheet, and it is one most assessment contracts and data governance policies were not written with in mind.

Generated items, real-time scores

Lovelace's claim on item quality deserves scrutiny rather than a nod. She says generative AI items "perform just as well" as ones built by the program's own subject matter experts, attributing that parity to the fact that expert-authored content was the material used to train the generation process. That is a meaningfully different claim than saying AI items match human-written items in general, and any buyer evaluating a similar tool should ask the same question: what source content trained the generator, and does that provenance hold for a test of your curriculum.

The scoring-speed claim is more straightforward and more consequential for workflow. "What AI has done is, it's been able to score, one, more quickly and, two, with more rich information," Lovelace said, describing results that previously took weeks or months arriving close to the moment of testing. For a district technology office, that compresses an entire reporting cycle that currently spans state assessment windows, vendor turnaround times, and intervention planning meetings into something closer to a single term.

From a score to a portrait

The most ambitious part of Lovelace's framing is aggregation. She describes AI pulling results across different subject-area systems to build a more complete "portrait" of a learner over time, one a superintendent or principal could use to see individual students alongside the patterns forming around them in a classroom, school, or district. "AI can grab information from multiple systems so the superintendent and principal are able to see" the broader picture, she said, describing a shift from isolated subject scores toward a connected view of how a student is developing across every system that touches them.

Lovelace frames the ambition in developmental terms too. AI, she said, can help educators get beyond compliance-driven accountability and "think about how to customize learning," using the connective tissue between subjects, critical thinking, collaboration, to inform what a teacher does next in the classroom. That framing is genuinely useful for instructional planning, built on a premise formative assessment theory has argued for years: the point of assessing a student is to change what happens in the next lesson, not just to produce a score that sits in a file.

Where the pitch becomes a data architecture problem

Building that cross-system portrait means integrating vendor platforms, standardizing data formats across subjects, and deciding who at the district or state level can see reasoning-level data about how an eight-year-old is thinking through a problem. None of that gets solved by picking a good assessment vendor. It gets solved by the same data governance discipline CIOs already apply to student information systems, extended now to a new category of behavioral and cognitive data that simply did not exist in structured, shareable form before generative AI made it possible to capture.

That is also where the stakes shift. A wrong answer on a bubble sheet is a data point about content knowledge. A transcript of a student's Socratic conversation with a chatbot, or an uploaded sample of their actual work, is a richer and more sensitive artifact, one that reveals how a child thinks rather than just what they know. District leaders evaluating these tools need to treat that distinction as a governance category of its own, with its own retention rules, access controls, and vendor contract language, rather than folding it into whatever data policy already covers test scores.

What this means for the next procurement cycle

Assessment has historically been one of the more conservative procurement categories in K-12, built around multi-year contracts, psychometric validation, and state accountability requirements that move slowly by design. Lovelace's framing suggests that conservatism is about to be tested by a technology moving considerably faster than the contracts built to govern it. A district that renews an assessment contract on the old rubric, speed, cost, standards alignment, risks missing the actual decision in front of it, which is whether the new system's data model and governance terms match what the vendor is now capable of collecting.

The practical move for any technology leader evaluating this category is to separate the pedagogical pitch from the data architecture question and demand answers on both. Ask what the AI was trained on, how scoring decisions can be audited, and precisely where reasoning-level student data is stored, aggregated, and shared once it leaves the individual assessment. Settling those answers while the contract is still being negotiated, while there is still leverage to demand specifics, gives a district real recourse. Waiting until renewal season leaves the governance terms exactly where the first vendor wrote them.

Tagged#news#edtech#education#learning#lms#ai-education#k-12#assessment#edtech-procurement#data-governance#generative-ai