AivexaNewsSearch
AI news for builders and product teamsChecked every hour

AI coding agents generate more code, but not more software

Collected Oct 9, 2026

A study by Harvard University researchers Fiona Chen and James Stratton examined how the adoption of AI coding assistants and agents changed engineering output across hundreds of companies, drawing on aggregated analytics from Jellyfish, which measures engineering team activity at a granular level. The dataset spans 300 million individual work events, such as commits and pull requests, plus issue management software data covering more than 700,000 employees at over 700 software development firms, from 2021 through March 2026.

To isolate the effect of AI tools, the researchers combined directly measured AI usage with analyses of GitHub activity to pinpoint when each company began introducing AI coding assistants, which mainly autocomplete human-authored code, or AI coding agents, which write and submit code autonomously from prompts. They then applied a difference-of-differences regression across variables measured before and after adoption at different times across organizations.

The raw output numbers are large and consistent. Introducing AI coding agents at a firm produced a 30 percent increase in total lines of code generated, a 20 percent rise in total commits, and a 23 percent increase in pull requests on average. Extra code, however, did not translate into extra software. The resolution rate for Issues and Epics tracked in tools such as Jira, which represent wholesale software features, showed no statistically significant change after AI tools arrived. The researchers also found no compositional shift in the size or complexity of those tracked issues across the AI introduction.

The explanation sits in code review. The average review process time, measured from a pull request being submitted to it being merged into the codebase, grew 49 percent on average after AI agents were introduced. Finer-grained indicators move the same way: the share of pull requests with changes requested nearly doubled, and the number of comments per pull request rose 35 percent.

Firms responded by shifting people toward review. The share of workers performing code reviews increased 14 percent after AI agents arrived. Looking at total active workers in the Jellyfish data and cross-referencing with LinkedIn data for those firms, the researchers state they cannot attribute significant employment changes to AI. That is a notable finding on its own: adoption of agents at scale, in this sample, did not show up as measurable headcount reduction.

AI has not yet taken over the reviewing either. By March 2026, 80 percent of the measured firms used some form of AI code review, but AI agents accounted for only 23.3 percent of all review comments and 10.8 percent of all pull requests. Humans were still responsible for the overwhelming majority of review work.

This suggests a fairly simple mechanism, and it is worth stating plainly as my own reading of the numbers. Agents lower the marginal cost of producing candidate code, so more of it arrives at the review gate. Review is a human judgment task with its own throughput ceiling, and it does not scale at the same rate. When the volume and heterogeneity of submitted changes rise, reviewers need more back-and-forth, more comments and more revision cycles, each of which adds latency. The gain shows up as activity, not as shipped features.

In practice, a 30 percent increase in lines of code with no change in issue or epic resolution is a warning about throughput metrics. Lines of code and commit counts are easy to instrument but weakly tied to delivered value, and they can rise for reasons that increase downstream cost. A likely trade-off for teams is that agent adoption redistributes effort toward review, testing and integration rather than removing it. If reviewers are already a constraint, adding agent-generated volume can lengthen cycle time even while the coding phase feels faster.

There are reasons to treat these results as a snapshot rather than a verdict. AI agents remain a relatively new part of the coding world, and their output has seen significant updates and upgrades even since the study's March 2026 data cutoff. By this point, 95 percent of firms in the study have implemented AI coding agents, and many are presumably still learning when and how to deploy them well. The study's authors note that coding-time versus review-time trade-offs could improve as teams gain experience with which problems suit an agent. Meanwhile, the low share of review comments and pull requests attributed to agents suggests automated review has plenty of room to grow.

Why it matters: engineering leaders evaluating coding agents now have firm-level evidence that faster code generation alone does not raise software output, which shifts the relevant metrics toward merged pull requests, review latency and feature completion. For developers, the practical implication is that review capacity, testing and integration discipline become the binding constraints, so teams that invest in smaller changes, clearer specifications and automated checks may capture more of the agent speedup than teams that simply enable agents everywhere. My inference is that tooling vendors and platform teams will face pressure to improve review automation and triage, since that is where the measured bottleneck now sits.

Read at Ars Technica · AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Study finds coding efficiency gains get "absorbed" by human review "bottleneck."