AI coding agents boosted code output without clear task-completion gains, study finds
Data from more than 700 software firms showed longer reviews and more requests for revisions. The researchers found no statistically significant improvement in tracked issue resolution.
A Harvard study using engineering activity data from more than 700 software firms found that AI agent adoption increased code output but did not significantly improve the rate at which tracked issues and larger features were resolved. The added proposals came with longer merge intervals and more review activity, suggesting teams may need to expand review capacity to realize gains from faster code generation. The observational results cover data only through March 2026 and do not establish how every team or newer tools perform.
01
After agent adoption, firms recorded 30% more lines of code, 20% more commits, and 23% more pull requests on average.
02
The average time from pull-request submission to merge rose 49%; comments per pull request increased 35%, and the share requiring changes nearly doubled.
03
By March 2026, 80% of measured firms used AI code review, while agents accounted for 23.3% of review comments and 10.8% of pull requests.
Software firms produced more code after adopting AI coding agents, but their tracked work did not show a statistically significant improvement in completion rates. A study by Harvard researchers Fiona Chen and James Stratton, covered by Ars Technica on October 9, found that review times rose sharply alongside the extra output.
More proposed changes, not clearly more finished work
The study distinguishes between coding assistants, which help complete code primarily written by people, and agents, which primarily write and submit code from prompts. Its agent-adoption results show a clear increase in activity: firms generated more lines of code, recorded more commits—saved code changes—and submitted more pull requests, or proposals to merge changes into a codebase.
The researchers also examined the resolution of issues and epics tracked in project-management tools such as Jira. These records capture tasks and larger software features, rather than simply counting code. After AI tools were introduced, their resolution rate did not change in a statistically significant way. The researchers found no corresponding shift in the size or complexity of the tracked issues.
That is a narrower finding than saying agents produce no additional software. The study detected gains in code-production measures, but not a statistically significant improvement in this measure of completed work. It does not establish that every team failed to benefit, or that every extra line of code was unnecessary.
Code production after agent adoption
+30%Lines of code
Average increase in lines of code generated after firms introduced AI coding agents.
+20%Commits
Average increase in recorded commits following agent introduction.
+23%Pull requests
Average increase in proposed code changes submitted for merging.
The extra work reached the review queue
The average interval between a pull request’s submission and its merge increased by 49% after agents were introduced. That measures the length of the review process, not necessarily hours a person spent actively reviewing. Still, the accompanying measures show more demands on reviewers: the share of pull requests requiring changes nearly doubled, and comments per pull request increased by 35%.
More workers also took part in checking code. The share performing reviews increased by 14% after agent introduction. Chen and Stratton interpret the results as a downstream bottleneck: efficiency gained while writing code was absorbed by the work needed to review it before it could enter the codebase.
AI review tools were already common in the sample. By March 2026, 80% of measured firms used some form of AI code review. Yet agents accounted for 23.3% of review comments and 10.8% of pull requests. Widespread adoption of review automation therefore coexisted with a process in which humans still handled most of the measured activity.
A large workplace sample, ending in March
The analysis used aggregated data from Jellyfish, which measures engineering-team activity. It covered 300 million work events, including commits and pull requests, across more than 700,000 employees at over 700 software-development firms. The records ran from 2021 through March 2026 and included data from issue-management software.
To identify when firms introduced assistants or agents, the researchers combined directly measured AI usage with analysis of GitHub activity. They then compared changes before and after adoption across organizations that adopted at different times. This observational approach examines workplace outcomes rather than a controlled exercise in which developers receive the same coding assignment.
The employment analysis likewise found no significant changes attributable to AI, using active-worker records and LinkedIn data. Both the output and employment findings describe the period observed, not a forecast of what newer agents will achieve. The March cutoff is material: the study does not measure results from tools or workflow changes introduced after that date.
Sources
arstechnica.comAI coding agents generate more code, but not more software
Reader comments
Newest comments first. Replies stay oldest first.