A Randomized Trial Finally Measured What AI Lesson Planning Actually Saves
AI & ML

A Randomized Trial Finally Measured What AI Lesson Planning Actually Saves

A UK government-commissioned RCT found Oak National Academy's free AI tool Aila cuts teacher lesson-planning time by 49 minutes a week, giving procurement teams the first vendor-independent benchmark for an AI education tool.

PublishedOctober 10, 2026
Read time6 min read
Share

A rare independent benchmark for an AI tool

Most claims about AI productivity gains in education arrive from the vendor that built the tool, wrapped in a survey of self-selected users. The trial behind Oak National Academy's Aila is different. The Education Endowment Foundation, a UK government funded research body, commissioned the National Foundation for Educational Research to run an independent randomized controlled trial across Year 3 to Year 6 teachers in England over ten weeks in autumn 2025. Teachers were randomly assigned to use Aila or to continue planning as normal, and independent assessors rated the quality of the lesson resources each group produced without knowing which group made them.

That design matters more than the headline number. Any enterprise buyer who has sat through a vendor pitch built on a 200 person satisfaction survey knows how little that evidence is worth once a tool meets a full deployment. EEF rated this trial's security as high, citing rigorous design, strong delivery, and confidence in the time saving finding. For a category where most evidence is self-reported and most vendors grade their own homework, a government funded RCT with a comparison group is the first result in education AI that resembles the kind of evidence a CFO would actually accept.

The number behind the number

The headline result is that teachers using Aila saved roughly 49 minutes a week on lesson and resource preparation. In absolute terms, the Aila group averaged about 2.5 hours a week on planning against 3.3 hours for the comparison group, a reduction to 76 percent of the baseline. Crucially, independent assessors found the resources produced under time pressure with Aila were no worse in quality than those built the old way. That combination, meaningfully less time with no measured quality tradeoff, is the exact shape of result that justifies a tool's existence rather than just its marketing.

But the trial also found real limits on impact. Only 81 percent of teachers assigned to the Aila group used it even once, it was used for about 7 percent of lessons in an average week, and usage declined as the term went on. Fewer than one in five teachers said Aila meaningfully changed their teaching or planning practice. The time savings were real and measurable. The transformation of classroom practice that vendors usually promise was not.

The CEO who undersold his own tool

What makes this trial unusually credible is how Oak's own leadership responded to it. John Roberts, Oak National Academy's co-founder and CEO, said publicly that the company's own early surveys had suggested bigger savings than the independent trial actually found. That is a rare moment in edtech marketing: a vendor confirming that internal data overstated a product's benefit relative to a third party's randomized result. Roberts also noted that the data was collected in late 2025, meaning the evidence already trails the pace at which the underlying AI models have improved.

Roberts described teachers adapting existing materials with Aila's help more often than asking it to generate whole lessons from scratch, a usage pattern that tracks with the modest time savings rather than a wholesale replacement of planning work. Oak is now folding Aila's capabilities more directly into its existing curriculum resources rather than treating it as a standalone chat assistant, a product decision that itself reads as an admission that the generate everything approach undersold what teachers actually wanted.

Where the time savings concentrated

The trial's subgroup findings are the part a workforce planning lead should read closely. The greatest time savings showed up among early career teachers, teachers with lower subject or pedagogical confidence, and teachers who were already more comfortable with technology. That is a familiar pattern from enterprise AI deployments generally: tools help newer or less confident staff close a skills gap faster than they help experienced staff do the same work marginally quicker. It is also the group with the highest attrition risk in most organizations, which gives the time savings a retention angle beyond pure productivity.

The trial also surfaced a context that any CIO evaluating a workload reduction tool should recognize. Almost all participating teachers said lesson planning was already stressful and that they spent too much time on it before the trial started. The tool did not need to be transformative to be worth deploying. It needed to measurably reduce a task that staff already found burdensome, and it did that, by a verified 23 percent, even if it did not change what teaching itself looks like.

The evidence bar this sets for every AI vendor

Set this trial against the typical AI procurement conversation inside a school system, a university, or any enterprise buying AI productivity tools for a large frontline workforce. Most vendor claims rest on pilot programs the vendor itself designed, surveys of whichever users opted in, or usage dashboards with no comparison group. EEF's trial gives buyers something different: a specific, independently verified number, 49 minutes a week, attached to a specific population, under specific conditions, with an honest accounting of where adoption fell off.

That is the evidence bar procurement teams should start writing into renewal contracts and RFPs for any AI tool claiming to save staff time, whether the staff are teachers, support agents, or claims adjusters. Ask the vendor for a trial design with a comparison group. Ask what percentage of licensed users actually used the tool weekly, not just what percentage logged in once. Ask whether the vendor's own internal data matched or overstated the independent result. Oak passed that last test by admitting it did not. Most vendors in this category have never been asked the question.

The roadmap implication

The lesson for any technology leader evaluating AI productivity tools at scale is not to expect transformation from version one. It is to expect a real, bounded, measurable reduction in a specific painful task, delivered to the people who need it most, with usage that decays unless the product keeps earning attention. That is a modest claim, and it is also a sustainable one, because it survives an independent audit rather than collapsing under it.

Oak's response, folding Aila's capability into existing workflows rather than running it as a separate assistant competing for attention, is the more interesting signal than the 49 minute figure itself. The durable AI deployments in any large workforce will look less like a new app employees have to remember to open and more like a capability quietly embedded in the tools people already use every day. Buyers evaluating their next AI contract should ask vendors to show that path, not just a demo of the standalone chatbot.

Tagged#news#edtech#education#learning#lms#ai-education#procurement#edtech-research#eef#nfer#oak-national-academy#randomized-controlled-trial#aila