A Randomized Trial of 464 Teachers Found an AI Lesson Planner Saves 49 Minutes a Week, and Barely Anyone Kept Using It
AI & ML

A Randomized Trial of 464 Teachers Found an AI Lesson Planner Saves 49 Minutes a Week, and Barely Anyone Kept Using It

The EEF's rigorous trial of Oak National Academy's Aila tool found real time savings and no quality loss, but usage declined over ten weeks because the tool was built for a workflow teachers do not actually follow.

PublishedOctober 6, 2026
Read time6 min read
Share

A trial designed to actually answer the question

Most claims about AI tools saving teacher time rest on vendor case studies or small-scale pilots that are difficult to generalize. The EEF's trial of Oak National Academy's Aila tool is a meaningful exception: a genuine randomized controlled trial covering 464 primary school teachers across 108 schools, run over a full ten-week term by NFER, with one group using Aila and a comparison group able to use other AI tools but not Aila specifically. That design isolates Aila's specific effect rather than measuring AI assistance broadly, and the EEF rated the resulting findings as high security, its top confidence rating for trial methodology.

The headline result from that rigorous design is genuinely positive on its own terms: Aila users spent 2.5 hours a week on lesson preparation and resource gathering, compared to 3 hours and 19 minutes for the control group, a reduction of roughly 49 minutes weekly, or about 24 percent. For a profession where workload and burnout are persistent, well-documented problems, a credibly measured 24 percent reduction in one major time category is a result worth taking seriously rather than dismissing as another unproven AI productivity claim.

The quality question this trial actually settles

A standard worry with any AI tool promising time savings on teaching materials is that the saved time comes at the cost of resource quality, a reasonable concern given how directly lesson plan quality affects student learning outcomes. This trial directly addresses that concern: independent assessors reviewing the resources produced by both groups found no measurable quality difference between Aila-generated materials and those the control group produced through their normal process.

That finding matters beyond this specific tool. It demonstrates that a well-designed AI lesson planning assistant can deliver real time savings without the quality tradeoff critics typically assume comes with AI-generated educational content, at least within the scope this trial measured. Schools and vendors evaluating similar tools now have a credible reference point showing that outcome is achievable, even if it is not guaranteed for every AI tool in this category.

Why usage declined despite the measured benefit

The trial's most operationally important finding is also its most counterintuitive: despite genuinely saving time without sacrificing quality, Aila usage declined over the ten-week trial period, and the tool ended up supporting only about 7 percent of lessons in a typical week even among teachers who had access to it. A tool that measurably works is still failing to achieve meaningful adoption, which is a different and in some ways more interesting problem than a tool simply not delivering value.

The trial's own explanation for that gap is specific and actionable: Aila was designed to generate lesson plans from scratch, while most teachers' actual day-to-day practice involves adapting and modifying existing materials they already have, rather than starting a new lesson plan from a blank page. That is a workflow mismatch, not a value mismatch, and it is the kind of gap that shows up clearly in a rigorous trial but would be easy to miss in a shorter or less structured pilot.

Who actually benefited, and what that implies for rollout strategy

The trial found that early-career teachers and those less confident in their subject knowledge got the most value from Aila, a group for whom generating a lesson plan from scratch is a genuinely harder task than it is for an experienced teacher who already has a library of materials to draw from and modify. Experienced teachers, by contrast, largely stuck with their own established process of adapting existing plans, since that process was already faster and more suited to their workflow than starting fresh with an AI tool.

That split argues for a more targeted rollout strategy than a blanket, school-wide deployment aimed at every teacher regardless of experience level. Schools considering tools in this category should prioritize early-career teacher onboarding and support use cases specifically, where the trial data shows the clearest benefit, rather than expecting uniform adoption and value across a staff with widely varying baseline planning workflows and needs.

What Oak and similar vendors should actually build next

The trial's clearest product lesson is that the next version of tools in this category needs to support material adaptation and editing as a first-class workflow, not an afterthought bolted onto a from-scratch generation tool. Teachers in this trial frequently needed to edit Aila's output before it was genuinely classroom-ready, and many ultimately preferred other AI tools specifically for ease of use in that adaptation and editing context, rather than continuing with a tool optimized for generation from a blank starting point.

That gap between what the tool was built to do and what teachers actually needed it to do is a design problem, not a market problem. There is a real, trial-validated demand for AI assistance with lesson preparation broadly, this trial proves that demand exists and that meeting it does not have to cost quality, the specific product shape just needs to match how teachers actually work rather than how a planning tool is conventionally built.

The broader lesson for school technology procurement

For school and district technology leaders evaluating AI tools generally, this trial is a useful model for the kind of evidence to demand before a wide rollout: a properly randomized trial measuring both quantitative outcomes, like time saved, and genuine adoption and usage patterns over a realistic multi-week period, carries far more weight than a vendor case study or a short pilot. A tool that performs well on a time-savings metric but poorly on sustained usage, as Aila did here, would look like an unqualified success in a shorter evaluation that only measured the former.

The practical procurement takeaway is to ask vendors directly for usage retention data over time, not just initial performance metrics from a short trial window, and to specifically probe whether the tool's design matches how staff actually work day to day rather than how the vendor assumes they should work. Aila's measured 24 percent time savings is real and worth taking seriously, but this trial also shows clearly that measured benefit alone does not guarantee the sustained adoption a school actually needs to realize that value at scale.

Tagged#news#edtech#education#learning#lms#ai-education#eef-trial#oak-national-academy#aila#teacher-workload#randomized-controlled-trial#lesson-planning-ai#edtech-efficacy