Microsoft's own data says CLI coding agents lift merged pull requests 24 percent
Cloud

Microsoft's own data says CLI coding agents lift merged pull requests 24 percent

The first large field study of Claude Code and Copilot CLI at Microsoft measured a real output gain, and revealed that adoption spreads through managers and peers, not mandates.

PublishedJuly 27, 2026
Read time7 min read
Share

The first real field study of a coding-agent rollout

On July 1, 2026, Microsoft researchers Emerson Murphy-Hill, Jenna Butler, and Alexandra Savelieva published a study of the company's own early-2026 rollout of command-line AI coding agents. The paper covers tens of thousands of engineers using Anthropic's Claude Code and GitHub's Copilot CLI, observed from January 5 to April 29, 2026. It is the first field study to use direct developer telemetry, rather than surveys or lab tasks, to measure both who adopts these tools and whether they produce enough output to justify their cost. For engineering leaders drowning in vendor benchmarks and anecdote, a rigorous look at a real enterprise rollout is overdue.

This approach matters because most numbers floating around AI coding come from controlled experiments or self-reported time savings, and both travel badly to a live organization. This study watches actual engineers, using the tools of their choice, inside a large production codebase, and applies a causal-inference method to separate the tools' effect from the background trend. The result carries more weight for a staff-plus audience precisely because it is messy, observational, and drawn from the same conditions your own teams work in. It grounds the productivity debate in evidence from one of the largest software organizations on earth, even if no single study can settle that debate for good.

The headline number, and its confidence interval

Adopters merged roughly 24 percent more pull requests than they would have without the tools, with a 95 percent confidence interval running from 14.5 to 33.7 percent. That is a large, statistically clear effect, and it held across the four-month observation window rather than fading after the novelty wore off. For leaders who have watched pilots produce a burst of enthusiasm and then flatten, the persistence is the more interesting finding. The lift did not evaporate, which suggests the CLI agents became part of how adopters work rather than a toy they abandoned once the trial dashboards stopped being watched.

The honest reading of the interval is that the true effect could be as modest as 14 percent or as strong as 34 percent, and either end changes the ROI math. A 14 percent throughput gain against heavy token spend is a different investment case than a 34 percent gain, and the study cannot tell you which your organization will land on. What it does establish is direction and durability: the effect is real, positive, and lasting in this population. Pair that with your own cost data, and you have the two halves of an ROI conversation that most teams have been having with only one number in hand.

Adoption is social contagion, not a mandate

The most actionable finding has nothing to do with output. First use spread through social exposure, meaning an engineer was far more likely to try the tools when peers, skip-level colleagues, and especially their direct manager were already using them. Adoption looked like contagion moving through the org chart, and prior experience with IDE-based Copilot also predicted who picked up the command-line agents. That pattern tells you the fastest way to spread a coding agent is to seed it with connected, respected engineers and with managers, then let proximity do the work. A top-down license grant with no social ignition leaves most of the potential adoption on the table.

This reframes rollout as a network problem rather than a procurement one. If you want penetration, you invest in visible internal champions, manager enablement, and forums where early adopters demonstrate real workflows to their teams. The corollary is a warning: adoption clustered by social ties can also leave whole pockets of the org untouched, not because those engineers are resistant but because no one nearby modeled the tool. Leaders measuring adoption by headline license counts will miss those gaps entirely. The study suggests tracking adoption by team and by reporting line, so you can see where the contagion stalled and intervene there deliberately.

Retention favors the already-active

The researchers defined an adopter as retained if they used the tool on at least 5 of the 14 days after first use, and found that sustained coding activity predicted retention more strongly than demographics or career stage. In plain terms, engineers who were already writing a lot of code kept the agents, while those who code less often drifted away. That is intuitive, but it has a sharp implication for where the investment pays back. The population that extracts the most value is the one already at the keyboard, which means the biggest returns concentrate among your most active builders rather than spreading evenly across every role.

For capacity planning, that concentration is useful. It argues for prioritizing seats and token budgets toward high-activity engineering teams first, where retention and output gains are most likely, before broadening to roles that touch code occasionally. It also complicates the common executive hope that AI coding will uniformly lift everyone. The study's dose-response analysis reinforces the point: more engagement correlated with more benefit. Leaders should therefore expect a skewed distribution of value, plan their rollout to feed the engineers who will actually run the agents daily, and resist the temptation to measure success by breadth of access alone.

The proxy problem the authors refuse to hide

To their credit, the researchers state plainly that merged pull requests are a proxy for output, not a measure of product value, maintainability, or code quality. A 24 percent lift in PRs could reflect genuinely more shipped work, or it could reflect the same work sliced into smaller commits, or faster generation of code that later needs rework. The telemetry cannot distinguish those cases, and the paper does not claim it can. Any leader citing the 24 percent figure should carry that caveat with it, because stripping the caveat turns careful research into the kind of inflated claim the study was designed to test.

This is where the study connects to the tooling conversation happening in parallel, as clouds ship dashboards that count the same commits and PRs. Those activity metrics are necessary but insufficient, and the Microsoft paper is a well-credentialed reminder of why. The durable question is whether faster PR throughput converts into faster, safer delivery, and answering it requires joining agent telemetry to change-failure rate, review latency, and incident data that no coding-agent dashboard captures on its own. The 24 percent is a real signal worth acting on. Treating it as the whole picture is the mistake the authors went out of their way to prevent.

What a staff-plus leader should take from it

The practical playbook writes itself. Seed adoption socially through managers and connected engineers rather than relying on a mass license drop. Concentrate early investment on high-activity teams where retention and output gains are strongest. Measure adoption by reporting line so you can find and fix the pockets the contagion missed. And instrument output beyond PR counts, joining agent data to your delivery and reliability metrics so the ROI case rests on outcomes rather than activity. None of these steps require new tooling, only the discipline to treat a coding-agent rollout as a change-management program with evidence behind it.

The broader signal is that the AI coding debate is maturing from marketing into measurement, and the numbers are landing in a credible middle. A real, durable 24 percent PR lift is neither the tenfold revolution some vendors sell nor the nothing-burger skeptics claim. It is a meaningful, costly, unevenly distributed gain that rewards deliberate rollout and honest measurement. Engineering leaders who internalize that framing will spend their token budgets where they pay back and defend their investment with data. Those still quoting headline productivity multiples will keep getting surprised when the plateau arrives, as this study quietly predicts it will.

Tagged#news#engineering#software-engineering#devops#platform-engineering#architecture#infrastructure#ai-coding-agents#claude-code#github-copilot#developer-productivity#microsoft-research#engineering-management#pull-requests#developer-experience#adoption#research