A Utah district has actual data on what AI tutoring does to critical thinking
AI & ML

A Utah district has actual data on what AI tutoring does to critical thinking

Jordan School District says two years of AI conversations across 82 teachers produced a 28 percent gain in critical thinking, real evidence at a moment when most AI-in-learning claims are still anecdotal.

PublishedSeptember 16, 2026
Read time6 min read
Share

What the district measured

Jordan School District in West Jordan, Utah, analyzed nearly 14,000 student to AI conversations collected from 82 teachers over two years and reports a 28 percent increase in student critical thinking skills, with higher level reasoning abilities more than doubling across subjects and grade levels over the same period. That is a specific, sourced number attached to a real data set, in a category where most public claims about AI's effect on learning remain anecdotal, self-reported by a single teacher, or lifted directly from a vendor's own marketing material.

Digital teaching and learning administrator Jared Covili described the motivation behind the rollout in modest terms: "We're always looking for innovative practices for students to improve their learning." The scale of the underlying data set, spanning two full years and dozens of participating teachers across multiple subjects, is what makes this worth genuine attention beyond the usual single classroom pilot story that circulates constantly in edtech vendor marketing and conference keynote slides.

The boundary that seems to matter

The district's design choice is the part enterprise buyers should study closely, because it is the piece most competing deployments skip. Covili explained the mechanism directly: "The AI can ask kids questions, not just the three or four kids that raise their hands." The tool's job in this deployment is prompting broader classroom participation and deeper, more probing questioning, not generating finished answers or content for students to submit as their own original work.

That boundary was enforced deliberately and communicated clearly to teachers before rollout. Covili was explicit about exactly where the district drew the line: "We don't want to turn over student writing to AI, or teacher interactions with students." AI functions here as a thought partner that asks questions back at the student, not a system that grades finished work or replaces the human teacher relationship, and that specific distinction appears to be doing real, measurable work in the outcome data the district collected.

Why this is an outlier finding right now

Most public discussion of AI in K-12 this school year has been about restriction, not results. New York City Public Schools and Los Angeles Unified, the two largest districts in the country by enrollment, both imposed yearlong moratoriums on student facing AI tools for 2026-27 rather than continue deploying and publish outcome data of their own the way Jordan District has now done. Jordan District's willingness to deploy at real scale and rigorously measure the result puts it in a genuinely small minority among large public school systems nationally.

That scarcity is exactly why the number deserves careful scrutiny rather than either uncritical adoption by other districts or blanket dismissal by skeptics. A district publishing a positive, quantified result while the two largest systems in the country are pausing entirely is a real signal about what deliberate scope limits and sustained teacher training can accomplish, though it is not proof that broader, faster deployment elsewhere would produce the same effect without the same two years of preparation and the same enforced constraints on what the tool is allowed to do.

The methodology questions enterprise buyers should still ask

A 28 percent gain over two years, measured through conversation analysis rather than a formally controlled study, invites reasonable questions before any enterprise buyer extrapolates it to their own context. What specific critical thinking instrument produced that number, was there a genuine comparison group of similar students without AI access to benchmark against, and how much of the measured gain reflects the intensive teacher training that accompanied the rollout rather than the AI tool itself doing the work.

None of that undercuts the finding's underlying value. If anything it sets the standard for what enterprise buyers should demand from their own AI-in-learning pilots going forward, corporate L&D programs very much included: a real, named evaluation instrument, a defined comparison baseline established before launch, and enough scale and duration in the pilot to separate a genuine, durable effect from novelty enthusiasm that reliably fades after the first semester or the first quarter of any new training rollout.

Two districts, two bets, one school year

Jordan District's bet and LAUSD's moratorium are not actually opposite positions on whether AI works in a classroom setting. They are different answers to a governance question that every large organization deploying AI eventually has to answer for itself: can we build and enforce the guardrails fast enough to deploy responsibly this year, or do we need a pause first to build them properly. Jordan District effectively says yes, backed by two full years of preparation and teacher training behind that answer. Los Angeles says not yet, applied city-wide, compressed into a single school year of planning time.

Both positions are defensible depending on an organization's actual governance readiness at the moment the decision gets made, and that is the real lesson enterprise leaders should take from the contrast. The right AI-in-learning decision is rarely deploy everywhere immediately or restrict everywhere indefinitely. It is matching the scope of deployment to the demonstrated maturity of your governance, the way Jordan District appears to have done with two years of structured teacher training completed before it published any result at all.

What this means for how you pilot AI in learning

Treat Jordan District's approach as a template for structure and sequencing, not as a guarantee of an identical outcome elsewhere. Define specifically and in writing what the AI is and is not permitted to do before any deployment begins, train staff thoroughly on those boundaries before scaling beyond a pilot group, and build a real measurement plan with a defined comparison baseline from day one rather than retrofitting an evaluation methodology once the pilot is already running and the easy baseline window has already closed.

The organizations that will have credible AI-in-learning ROI data a year from now are the ones setting up that measurement infrastructure today, whether the deployment in question is a K-12 classroom or an enterprise corporate training platform serving thousands of employees. Jordan District's number is useful precisely because it exists at all, backed by two years of data and a named methodology. Most competing claims circulating in this category right now still do not have anything comparable behind them.

Tagged#news#edtech#education#learning#lms#ai-education#jordan-school-district#ai-tutoring#learning-outcomes-roi#k-12-ai-deployment