A rare real trial in a sea of assumptions
Most universities that roll out a generative AI tool do it on faith, a vendor pitch deck, and a pilot group that volunteers itself into a biased sample. The University of Maryland did something different. Before offering its Virtual Study Assistant campuswide, it asked a team led by Jing Liu, director of the university's Center for Educational Data Science and Innovation, to run an actual randomized controlled trial. During fall 2025, 2,379 undergraduates and 30 instructors across multiple disciplines were split into VSA users and non-users, and researchers compared their course participation and grades.
That is the study higher ed critics have been asking for since ChatGPT arrived on campuses nearly four years ago, and almost nobody has run it. Provost Jennifer King Rice told Inside Higher Ed the results were not what she hoped. "We don't have clear findings from the study. It raises questions that haven't yet been answered," she said, adding that the university needs to understand what support is needed "rather than simply turning this on." That restraint, choosing to withhold a campus-wide launch pending more evidence, is itself the more interesting story for any CIO watching the AI procurement cycle.
What the trial actually found
The topline result is uncomfortable for anyone who assumed access alone drives adoption. Among the students randomly given VSA access, only about 15 percent chose to use it at all. Of that smaller group, 73.8 percent used it to get information, explanations, or direct solutions, 11.3 percent used it for practice and test preparation, and just 0.7 percent asked it to critique their own work. The study also found that VSA access correlated with marginally lower final grades and LMS participation across sections of the same course, a finding the researchers treat as a flag rather than a verdict.
Liu's team is careful about causation, and so are we. Students in the trial still had access to ChatGPT and Gemini through other university channels, which muddies any clean comparison. "You cannot force people to only use one tool," Liu said. Justin Reich, who directs MIT's Teaching Systems Lab and was not involved in the study, put it bluntly: a randomized trial in the field is "a bundle of interventions," and researchers "can't precisely tease out all of the causal mechanisms." For CIOs, that is a reminder that a single study rarely settles a procurement question on its own.
The default mode problem
The most actionable finding in the study has nothing to do with the model and everything to do with configuration. VSA ships with two modes: a default that returns the correct answer immediately, and a tutor mode that, per the study, "guides students toward a correct response." Few instructors ever switched their sections into tutor mode. That single setting, left on default, likely explains much of why access did not translate into learning gains, and it is the kind of detail that gets lost between a vendor demo and a faculty rollout.
Liu was direct about where responsibility sits. "Oftentimes we just throw AI tools to students, but instructors play such an important role in helping students to learn about these tools and how to use them effectively," he said, noting that faculty received only "light training" before the trial began. For any institution deploying an AI tutor, lesson planner, or study assistant, this is the governance question that matters more than the model choice: who owns the default configuration, and who is accountable for checking it against the instructional intent.
Why building your own rarely wins
Maryland's VSA was built in-house, wired directly into the campus LMS, and designed to draw only from each course's own materials, a reasonable response to data governance concerns about sending student work to third-party tools. But Stephen Aguilar, an associate professor of education at USC who was not involved in the study, thinks the low usage reflects something structural. "Traditionally, educational institutions are bad at creating bespoke or modified ed-tech tools for their populations. We can't compete against the market," he said.
Aguilar's suggested alternative is narrower, not grander: smaller, more personalized models that do a couple of things well, rather than a general-purpose assistant competing against ChatGPT and Gemini for the same fifteen minutes of a student's attention. That is a useful frame for any enterprise buyer evaluating an internal build. The question is rarely whether your engineering team can produce a working AI tool. It is whether that tool can win against free, familiar, general-purpose alternatives your users already have open in another tab.
The governance takeaway for every CIO
Set aside the grade data, which the researchers themselves do not treat as conclusive. The durable lesson from Maryland's trial is procedural: it ran a controlled study before a full rollout, found an uncomfortable result, and used that result to delay rather than accelerate deployment. That sequence, pilot, measure, decide, is exactly what most AI governance frameworks claim to require and what most actual rollouts skip under pressure to show momentum.
If you are weighing an internal build against a vendor contract for tutoring, coaching, or any other learning tool, Maryland's experience argues for sequencing the measurement step ahead of the adoption step, running the controlled pilot while the stakes of a wrong configuration are still small. It also argues for auditing configuration defaults as rigorously as you audit the underlying model, since the setting an administrator never checked can matter more than which foundation model sits behind the interface. Reich's caution about field trials being "a bundle of interventions" should push every team toward better instrumentation before committing to a campus-wide or company-wide scale-up, because the alternative is discovering the same uncomfortable gap only after thousands of users depend on the tool working as advertised.



