Coding Agents: 6 Signals It’s Not the Tool You Think

September 29, 2026

Adoption of AI coding agents nearly doubled in the first half of 2026, and developers say Claude Code is the tool they love most. Microsoft’s own field study of its engineers found something else entirely: the tool that produced the bigger measured lift in shipped code wasn’t the one anyone expected. Six signals from primary research on what coding agents are actually doing to engineering teams, not what the surveys about them feel like.

Stats band showing four key figures the piece is built on: JetBrains 2026 finding 39% of developers now run Claude Code, up from 18%; Microsoft's causal study finding a 24% lift in merged code; LinearB finding AI-assisted pull requests wait 5x longer for review; and Fortune reporting one engineer's monthly token usage cost Meta $1.4M

Six months into 2026, the story about coding agents seemed settled: adoption was climbing fast, developers had a clear favorite, and the productivity case was closed. Anthropic’s Claude Code went from a niche coding agent to the most-used one in the industry in under a year, and developer surveys kept ranking it the most loved. That’s the story most teams walked into their own coding-agents rollout with.

Then Microsoft ran the study that actually mattered: not another survey asking engineers how productive they felt, but a field study of tens of thousands of its own engineers, measuring what they actually shipped. The results complicate almost every assumption in the popular version of this story.

Coding agents are producing a real, durable lift in merged code, but the coding agent developers say they love isn’t the one the data says performs best, the productivity gain from coding agents isn’t evenly distributed across a team, and the bottleneck coding agents create shows up somewhere nobody was measuring. Six signals from primary research on what coding agents actually do inside a real engineering organization, not from another survey about how using coding agents feels.

Adoption climbed while trust fell

Card showing Claude Code adoption reached 39% of developers in mid-2026, up from 18% six months earlier, while Copilot and Cursor usage both declined over the same period

Two numbers that should move together, adoption and trust, moved in opposite directions for coding agents in 2026. That divergence between how many developers use coding agents and how much they trust them is the starting point for almost everything else in this piece.

Claude Code adoption nearly doubled in six months

JetBrains’s State of Developer Ecosystem 2026 survey, fielded across more than 15,000 professional developers between May and July, found Claude Code usage at 39% globally, up from 18% in January, with US adoption reaching 47%. GitHub Copilot moved the opposite way, down to 21% from 29% a year earlier, and Cursor slipped from 18% to 12%. 

GitHub’s own Octoverse 2025 report tells a similar story from inside its own platform: more than a million pull requests were opened by a coding agent between May and September 2025 alone, and nearly 80% of new developers on the platform started using Copilot within their first week of signing up. Whatever else is true about coding agents in 2026, developers are not adopting coding agents at a steady, predictable pace. They are moving toward one tool in particular, fast, and reaching for some form of coding agent almost as soon as they get access to a codebase at all.

Stack Overflow’s own survey found trust moving the other way

Stack Overflow’s 2025 Developer Survey, based on more than 49,000 responses, found that trust in AI accuracy fell to 29%, down from 40% the year before, while overall favorability dropped from 72% to 60%. Forty-five percent of respondents named output that’s “almost right, but not quite” their single biggest frustration with coding agents, and 66% said fixing that almost-right code cost them more time than writing it from scratch would have. 

Usage and confidence are not the same metric, and for coding agents in 2026 they told two different stories at once. A team reading only the adoption numbers would conclude coding agents had won the argument; a team reading only the trust numbers would conclude the opposite, and both would be looking at the same population of developers in the same year.

The productivity lift is real, and it doesn’t fade

Card showing Microsoft's own engineers merged 24% more code using coding agents, with the productivity lift holding steady across four consecutive months with no measurable decay

Every question about coding agents eventually comes back to whether they actually produce more shipped work, not just a feeling of speed. Microsoft’s own engineers, using coding agents daily inside a real engineering organization, answered that question with the best evidence available: what they actually merged.

Microsoft’s engineers merged 24% more code, with a rigorous causal design

Microsoft’s own field study of its command-line coding agents, led by researchers Emerson Murphy-Hill, Jenna Butler, and Alexandra Savelieva and covering tens of thousands of engineers, used a Bayesian synthetic-control design to compare adopters against a modeled counterfactual of how much they would have merged without the tools. The result: a 24.0% lift in merged pull requests per engineer per day, with a 95% credible interval of +14.5% to +33.7% and a posterior tail-area p-value below 0.001. 

That is not a self-reported productivity feeling. It is a measured increase in the actual unit of output coding agents are supposed to move, built from a synthetic counterfactual constructed out of ten independent control groups of engineers who never touched either tool, which is a materially harder bar to clear than the before-and-after comparisons most coding-agent productivity claims rest on.

The lift held steady for four straight months, unlike a prior study of Cursor

Earlier research on Cursor adoption in open-source projects found an early productivity lift that faded within a few months as novelty wore off. Microsoft’s researchers tested for the same pattern and did not find it: the lift measured +29.4% in February versus +20.0% across March and April, with overlapping confidence intervals that indicate no statistically meaningful decay. 

One surveyed engineer put it this way: “I no longer think about narrow solutions; instead I am able to use agents to think broadly and formulate wholistic approaches. I feel so much more productive, I am never going back.” For coding agents specifically, in this one large, well-instrumented rollout, the gain looks durable rather than a honeymoon effect. That distinction matters for how a rollout gets budgeted: a benefit that fades in three months is a pilot expense, while a benefit that holds for four and counting is closer to a standing capability a team should plan its headcount and tooling spend around.

The tool developers love isn’t the tool that performs

Bar chart comparing measured pull request lift between Copilot CLI at 24.9% and Claude Code at 11.4%, next to developer favorability data showing Claude Code far more loved despite the smaller measured lift

This is the finding that undercuts the popular narrative about coding agents most directly. The same Microsoft study that measured the 24% average lift also compared the two coding agents its engineers actually used, and the result runs against what every adoption survey about coding agents would predict.

Copilot CLI adopters saw roughly twice the lift Claude Code adopters did

Comparing each engineer’s own weeks of tool use against their own zero-tool weeks, Microsoft’s researchers found Copilot CLI adopters gained +24.9% in merged pull requests, while Claude Code adopters gained +11.4%, a difference of roughly 2.2 times, significant at p < 0.0001. The study’s authors are candid that they cannot fully explain why: their internal survey only hints at “shifting perceptions,” with some engineers describing both tools as “impressively good and only getting better” while migrating toward Copilot CLI regardless. 

Whatever is driving the gap, it is not visible in developer sentiment alone, and it means the two most common ways teams currently pick a coding agent, asking engineers what they prefer, or defaulting to whichever tool is already bundled into an existing platform subscription, are both answering a different question than the one that actually determines shipped output.

Claude Code is still the one developers reach for, and love most

That performance gap sits awkwardly next to coding agents adoption data. JetBrains’s same 2026 survey found Claude Code the most loved coding tool among developers at 46% favorability, far ahead of Cursor at 19% and Copilot at just 9%, even as Copilot held more total workplace adoption at 29% globally that same period. 

A team choosing a coding agent based on which one developers say they’d recommend, and a team choosing based on which one Microsoft’s own causal data says lifts merged output more, could reasonably land on different tools. That gap between preference and performance is exactly the kind of thing a survey alone will never surface, and it’s the strongest argument in this entire body of research for measuring a coding-agent rollout against actual merged output rather than a satisfaction score collected a few weeks after launch.

The bottleneck moved from writing code to reviewing it

Card showing AI-assisted pull requests wait over 16 hours for reviewer pickup versus roughly 200 minutes for unassisted work, with a 30-day merge rate of 32.7% versus 84.5%

If coding agents genuinely produce more shipped code, the obvious next question is what happens downstream of that code, once it reaches another human. The answer, at scale, is a bottleneck nobody building a coding-agents rollout plan was measuring a year ago.

AI-assisted pull requests are bigger and wait far longer for a reviewer

LinearB’s analysis of 8.1 million pull requests, drawn from 4,800 teams across 42 countries, found AI-assisted pull requests running more than 400 lines at the 75th percentile against 157 lines for unassisted work, roughly two and a half times larger. Those larger AI-assisted pull requests then waited more than 16 hours on average before a reviewer picked them up, against roughly 200 minutes for unassisted work, a delay more than five times longer. 

The code coding agents produce doesn’t just get written faster. It gets queued slower, and a review queue that grows faster than the team assigned to clear it is invisible in every metric that only tracks how quickly code gets written in the first place.

Once a reviewer starts, AI-assisted code actually moves faster, but less of it ships within 30 days

The nuance LinearB’s data surfaces is that the review itself, once started, runs faster for AI-assisted pull requests at roughly 194 minutes against 252 minutes for manual work. The real cost sits entirely in the wait, not the read. That wait shows up in the outcome that matters most: only 32.7% of AI-assisted pull requests merged within 30 days, against 84.5% for unassisted work. Coding agents didn’t create a review problem in the traditional sense. 

They created a queueing problem that happens to look, from the outside, like a review problem. There’s a partial answer already circulating inside the same ecosystem: GitHub’s Octoverse data found 72.6% of developers reporting improved review effectiveness when they used Copilot itself as part of the review step, which suggests the fix for a coding-agent-driven backlog may be a second coding agent working the queue from the review side, not fewer agents writing code in the first place.

Junior engineers and senior managers gain the most

Bar chart showing pull request lift by career stage forming a C-shape, with junior individual contributors at +44% and senior managers at +34%, both larger than the +21.3% baseline for mid-level contributors

Not every engineer benefits from coding agents equally, and the shape of that variation is not what a flat, one-size-fits-all coding-agents rollout plan would assume.

The lift follows a “C” shape across career stages

Microsoft’s within-person analysis, evaluated at three tool-use days per week, found a mid-level individual contributor baseline lift of +21.3%. More junior individual contributors and more senior managers both showed larger gains, up to +44% at the most junior level measured and +34% among the most senior managers, while engineers at the mid-career reference point saw the smallest lift of the group. Coding agents appear to help most at the two ends of a career ladder, not the middle, which cuts against the common assumption that coding agents are mainly a junior-engineer productivity tool.

Tenure shows the same curve, with one result the researchers flagged as uncertain

The same pattern repeated by tenure at the company: engineers with 5 to 15 years at Microsoft, the reference group, saw a +21.5% lift, while engineers with 15 or more years saw +34%. Engineers with less than a year of tenure showed the largest point estimate in the entire study, though the researchers explicitly cautioned that this figure is likely conflated with ordinary onboarding ramp-up and should be read with skepticism even after their attempts to control for it. 

The honest, unresolved edge of the data is itself useful: not every large number a coding-agents study produces should be taken at face value, including some of the ones inside the same study, and a team rolling out coding agents to a mixed-tenure engineering org should expect the gains to land unevenly rather than plan around a single average.

The cost side is no longer hypothetical

Card showing Meta's highest individual AI token user consumed 281 billion tokens in one month, worth over $1.4 million at list price, with total company usage exceeding 60 trillion tokens in the same window

None of the productivity case for coding agents matters to a budget owner if the cost side stays invisible. In 2026, for at least one organization running coding agents at scale, it stopped being invisible in a very public way.

One engineer’s usage alone could run into seven figures

Fortune’s reporting on Meta’s internal AI usage dashboard found that in a single 30-day period, Meta’s highest-ranked individual token user averaged 281 billion tokens. At the lowest published price for Claude Opus, $5 per million tokens, that one engineer’s usage alone could have cost Meta more than $1.4 million in a single month. 

Meta’s own CTO Andrew Bosworth, discussing a top-performing engineer whose coding-agent usage cost roughly matched their salary, reportedly waved off the concern: “It’s like, this is easy money. Keep doing it. No limit.” Not every organization running coding agents at scale will have Meta’s balance sheet to absorb that kind of spend without a second thought.

The token economics are now an industry-wide line item, not an edge case

Total employee usage of coding agents and other AI tools on that same Meta dashboard exceeded 60 trillion tokens in the 30-day window, which at Anthropic’s list pricing would run to roughly $900 million a month if paid at full retail rates. 

Meta is an unusually large, unusually well-resourced case, but the shape of the problem generalizes: once coding agents move from pilot to default tooling across an engineering organization, the token bill stops being a rounding error on the cloud invoice and starts being a budget line that needs its own governance, the same way seat licenses or cloud compute already do.

What this means for how a team actually adopts coding agents

Six signals, and none of them point toward the simple version of the coding agents story. Adoption is real and accelerating. The productivity lift is real and durable. But the coding agent developers say they love isn’t necessarily the one the data says performs best, the gains concentrate unevenly across a team, the bottleneck that used to be about writing code has quietly become a queueing problem in review, and the cost side has stopped being theoretical.

That combination argues against copying whatever coding agent a team’s engineers already prefer and calling a coding-agents rollout done. It argues for measuring merged output the way Microsoft did, watching the review queue the way LinearB’s data suggests, and building in real capacity, not just tooling, to absorb the load once code starts arriving faster than reviewers can clear it. 

That capacity question is exactly where a lot of engineering organizations get stuck: the code review bottleneck LinearB measured doesn’t resolve itself just because a team adds another coding agent license. It resolves when there are enough of the right engineers, with the right context on how the team’s own coding agents are configured, to actually clear the queue.

Landskill’s own Developer Productivity piece covers the broader measurement problem this connects to, and the RAG pipeline piece looks at a related class of AI infrastructure decision engineering teams are making under similar pressure. For teams whose gap is reviewer capacity rather than agent licenses, IT Outsourcing and Nearshore exist to add exactly that kind of engineering capacity on a timeline shorter than a permanent hire. Get in touch to talk through where the queue is actually breaking on your team.

What´s New

Deployment frequency looks great on a dashboard and tells leadership almost nothing on its own. This piece breaks down six developer productivity signals from 2025 and 2026 research - DORA, DX, and GitHub's own platform data - that actually hold up once you dig past the headline number.
Most RAG failures trace back to the vector database layer, not the model. This guide breaks down six production decisions — from index choice to chunking strategy to observability — that determine whether a RAG pipeline actually retrieves the right context.
WebAssembly crossed a real threshold in 2025 and 2026: version 3.0 shipped, WASI's component model matured, and JVM languages gained working paths to target it. This piece breaks down six concrete shifts, from edge production numbers to where Wasm still isn't ready, so backend and platform teams can decide where to invest now.
We value your privacy

We use cookies to run this website, measure its performance and show relevant content. Cookies Policy