The number worth pausing on
Jordan School District in West Jordan, Utah adopted an AI tool roughly two years ago and has been tracking its effect on student reasoning ever since. The district's analysis, drawn from nearly 14,000 student-to-AI conversations across 82 teachers, found a 28 percent increase in critical thinking skills and more than a doubling of higher-level reasoning abilities across subjects and grade levels. That is a rare thing in this space: a district reporting a specific, tracked outcome measure rather than an adoption rate, a satisfaction score, or a usage count.
The contrast with the broader national mood matters. Surveys elsewhere this fall show a majority of teachers believe AI is harming students' critical thinking, which is part of what is driving districts like New York City and Los Angeles toward outright pauses. Jordan is reporting the opposite result from the same category of tool. The obvious question is what the district did differently, and the answer is not the technology. It is how the district chose to constrain it.
The governance choice underneath the result
Jared Covili, the district's digital teaching and learning administrator, described the operating principle plainly: the AI is positioned as a thought partner, not a content creator. "The AI can ask kids questions," Covili said, "not just the three or four kids that raise their hands." That is a specific design constraint, not a marketing description. The district deliberately did not turn the tool loose to write for students or to mediate teacher-student interaction directly. Covili was explicit about the boundary: "We don't want to turn over student writing to AI, or teacher interactions with students. You don't want to take the human out of the loop."
That constraint is the entire mechanism behind the 28 percent figure. A tool that answers questions for students removes the reasoning step that produces critical thinking gains. A tool that asks questions back, at a scale no single teacher managing thirty students can match individually, adds practice reps in reasoning that would otherwise go to a handful of students willing to raise their hands. The result is not evidence that AI improves critical thinking in general. It is evidence that this specific governance choice, question-asking over answer-giving, produces a measurable effect when applied consistently for two years.
Why this is a governance story, not a product story
Every major AI vendor selling into education can technically support a thought-partner mode, an answer-giving mode, or both, often inside the same product. The differentiator in Jordan's result is not which vendor's tool sits underneath the deployment. It is that the district defined the acceptable use case before rollout, trained 82 teachers to hold that line, and then measured the outcome over two full years rather than declaring victory after a semester of adoption data. Most K-12 AI rollouts skip at least one of those three steps, usually the two-year measurement, because the pressure to show fast usage numbers outweighs the patience required to show learning outcomes.
This maps directly onto the pilot-to-production problem enterprise technology leaders already recognize from their own AI deployments. A copilot tool that drafts an employee's analysis for them and a copilot tool that questions the employee's draft until it holds up are the same underlying model wearing two different governance policies, and they produce very different long-run effects on the workforce's own reasoning capacity. Jordan's data is a rare, multi-year natural experiment showing that the governance policy, not the model choice, is what determines whether the deployment builds capability or erodes it.
The limits of a single district's number
One district's self-reported figure deserves the same scrutiny any technology leader would apply to a vendor case study. Jordan School District has not published its methodology in a peer-reviewed venue, the 28 percent figure comes from the district's own analysis, and 82 teachers across one Utah district is a modest sample against the scale of national K-12 enrollment. None of that makes the result meaningless, but it means the honest reading is a promising internal case study, not a proven, generalizable formula ready to copy into a different district or a different vendor's tool without adaptation.
What does generalize is the discipline behind the number rather than the number itself: define the constraint before rollout, hold teachers to it for years rather than months, and measure the specific capability you actually want to protect. Any organization citing Jordan's result as justification for its own AI program should be able to answer whether it has adopted that same discipline, not just the same category of tool, because the discipline is what produced the result and the tool alone almost certainly would not have.
What to steal from a school district's playbook
The transferable lesson for any enterprise AI rollout is Jordan's sequencing: define the specific behavior you want the tool to reinforce before deployment, not after adoption metrics come back strong, and pick a mode, question-asking, draft-checking, gap-flagging, that keeps a human generating the actual output rather than merely approving the AI's. Then commit to measuring the capability outcome you actually care about, not just usage volume, on a timeline long enough to see it, which in Jordan's case meant two years and nearly 14,000 logged interactions before publishing a number.
For a CTO or CIO weighing whether an AI pilot in your own organization is building capability or quietly substituting for it, Jordan's model gives you a concrete test to apply: can you point to a specific design decision, documented before rollout, that explains why your deployment should build reasoning rather than replace it, and do you have a measurement plan running long enough to prove it either way. If the honest answer is that you are tracking adoption and satisfaction but not the underlying skill, you do not yet know which outcome you are actually getting, and neither will the people eventually asked to defend the program.



