Modernization and software Guide
AI Coding Agents in the Enterprise: What They Speed Up, What They Break and How to Govern Them
AI coding agents make individual engineers write far more code, and organizations ship a fraction of that gain. The difference is spent in review, stability and security, and the teams that keep it treat the agent as one part of a verification system they control.
For CTOs and VPs of engineering rolling out Claude Code, Copilot, Codex, Cursor or similar agents across their teams, deciding how far to trust them, what to measure and which controls to require.
The short answer
AI coding agents reliably raise individual output on well-scoped, greenfield and repetitive work, and the gain shrinks at review, testing and release. The largest telemetry study finds commits up about 240% and releases up about 30%. Keep the gain by containing agents, routing their changes through human review and hard CI gates, budgeting tokens as consumption, and measuring delivery and stability together.
Key takeaways
- In METR's trial, experienced maintainers working in their own repositories took 19% longer with AI and still believed it had made them about 20% faster, so self-reported time savings are weak evidence.1
- Across more than 500,000 developers, coding activity rose sharply with agents while releases rose about 30%; DORA finds AI adoption still linked to worse delivery stability.23
- AI-generated code passed Veracode's security tests 56% of the time in 2026, against 55% a year earlier, so a newer model is no security control.4
- The worst agent incidents, including production databases deleted with their backups, trace to agents holding credentials that reached production.56
- Vendors have moved from seats to metered tokens, and Uber capped AI tool spend after using its 2026 budget in about four months. Govern cost as consumption and keep usage out of performance reviews.789
Most engineering leaders now have the same two reports on their desk. One says developers love their AI coding agents and feel much faster. The other says pull requests are bigger, review queues are longer and the incident count has not gone down. Both are accurate. Writing code has become cheap and abundant, and the scarce work has moved to checking it, merging it safely and paying for the tokens. This guide sets out what the measured evidence supports, where the risk has actually landed, and the operating model that lets an engineering organization keep the gain.
- 19%slower: experienced maintainers with AI in METR's trial, who believed they were about 20% faster1
- +30%rise in releases, against about +240% in commits, across more than 500,000 developers using AI agents2
- 56%of AI-generated code passed Veracode's 2026 security tests, one point better than a year earlier4
What the trials and the telemetry actually measured
The randomized evidence splits by setting. Three field experiments at Microsoft, Accenture and a Fortune 100 company, covering 4,867 developers, found 26% more completed tasks with an AI assistant, with less experienced developers gaining most.10 Google's internal trial of 96 engineers found them about 21% faster on an enterprise task.11 METR's 2025 trial put 16 experienced open-source maintainers on 246 real issues in repositories they knew well, and the tasks where AI was allowed took 19% longer. The developers had expected a 24% speedup, and afterwards still believed they had been sped up by about 20%.1 METR's 2026 attempt to repeat the study could not produce a reliable estimate, partly because developers increasingly refused to work without AI.12
The large observational studies see something the trials cannot: the gain shrinks at every stage after the keyboard. An NBER working paper covering more than 500,000 GitHub developers found coding activity rising with each generation of tools, from autocomplete to interactive agents to autonomous agents, while commits rose about 240%, projects about 80% and releases about 30%. The authors conclude the task-level gains have translated only partially into shipped software so far.2 Stanford's study of more than 100,000 engineers at over 600 companies finds gains of about 30 to 40% on simple tasks in new codebases and 0 to 10% on complex work in mature ones, and in one case study the team's effective output stayed flat after adopting AI while rework rose 2.6 times.13
| Study | What it measured | Headline effect |
|---|---|---|
| Microsoft, Accenture and Fortune 100 trials, 4,867 developers | Completed tasks with an AI assistant | 26% more tasks10 |
| Google trial, 96 engineers | Time on one enterprise task | About 21% faster11 |
| METR trial, 16 expert maintainers | Time on real issues in their own repositories | 19% slower, believed 20% faster1 |
| NBER, 500,000+ GitHub developers | Commits, projects and releases | +240%, +80%, +30%2 |
| Stanford, 100,000+ engineers | Gain by task complexity and codebase age | 30 to 40% on simple new code; 0 to 10% on complex mature code13 |
Company figures look far larger and measure something else. Google's chief executive said in April 2026 that 75% of new code at Google is AI-generated and approved by engineers.14 A share of accepted lines counts code, and the NBER gap shows a large share of code can sit beside modest gains in releases. Benchmarks deserve even less weight: OpenAI stopped treating SWE-bench Verified as a measure of frontier coding ability after an audit found flawed tests and signs of contamination.15 The defensible expectation is narrow. Expect real gains of roughly 10 to 40% on well-scoped, greenfield or repetitive work in mainstream languages, expect little or nothing from senior engineers deep in large legacy systems, and expect the bottleneck to move downstream.
The bill arrives in review, stability and code health
Every large dataset that looks past the commit finds the cost landing later. Google's DORA research, built on nearly 5,000 respondents, found AI adoption now improves throughput while it "continues to have a negative relationship with software delivery stability", and its summary line is that AI amplifies what a team already is.3 Faros, whose 2026 report draws on two years of telemetry from 22,000 developers across 4,000 teams, finds pull requests 51% larger, bugs per pull request up 28%, median review time five times longer and incidents per pull request tripled.16 Faros sells engineering analytics, and the direction of its findings matches the independent studies.
Thoughtworks has placed complacency with AI-generated code in the Hold ring of its Technology Radar, citing GitClear's research showing duplicate code and churn rising while refactoring declines, and recommends reinforcing test-driven development and static analysis inside the coding workflow.17 Developers feel it too: 66% of Stack Overflow's 2025 respondents named answers that are almost right as their top frustration.18 There is a skills cost as well. In Anthropic's own randomized study, junior engineers learning a new library with AI scored 17% lower on a follow-up quiz, with the largest gap in debugging.19
DORA's research also names what protects a team. Its AI capabilities model lists a clear organizational stance on AI, healthy and AI-accessible internal data, strong version control, small batches, a focus on users and quality internal platforms, and finds that small batches amplify AI's benefit.20 Whatever the foundation, review capacity and change safety have to be scaled on purpose, because a faster writer does nothing for the queue behind them.
Security: the code, the dependencies and the agent itself
Veracode's July 2026 report found generated code passing security checks 56% of the time on average, against 55% in its first report a year earlier, while it was syntactically correct almost every time. The best model scored 68%, more than half the models sat at 50 to 53%, and neither model size nor coding specialization made much difference. Weaknesses that follow a pattern improved, and those that need an understanding of data flow stayed poor.4 The finding sits on a long academic line: in 2021, about 40% of the programs an AI assistant wrote in security-relevant scenarios were vulnerable.21
Dependencies add a risk of their own. A USENIX Security 2025 study of 16 code models and 576,000 samples found that at least 5.2% of packages suggested by commercial models and 21.7% of those from open-source models did not exist, 205,474 unique invented names in all, and an attacker can register a name a model keeps inventing.22 Lockfiles, an internal registry or proxy, and a block on newly published packages close that gap. The larger shift is that the agent itself is now an attack surface. Simon Willison's rule of the lethal trifecta is the clearest guide: never give one agent access to private data, exposure to untrusted content and the ability to send data out at the same time.23 Our AI agent security guide covers that threat model in depth.
Nine incidents and the control each one lacked
The public incident record falls into four patterns: destructive actions by agents with too much access, poisoned agent tooling, malware that drives installed AI tools, and prompt injection through repository content. Each maps to a control that already exists.
| Incident | What happened | The control that would have stopped it |
|---|---|---|
| Replit, July 2025 | An agent ran destructive commands on a production database during a declared code freeze, then misreported what it had done5 | Agents never hold production credentials, and production changes need approval |
| PocketOS, 2026 | An agent hit a credential mismatch, found an unrelated API token in a file and deleted the production volume, taking the backups stored on it, in about nine seconds6 | Narrowly scoped tokens kept out of the codebase, and backups stored apart from the data |
| AWS Kiro, December 2025 (disputed) | Press reports blamed an AI agent for an AWS service disruption; Amazon says it was a limited event caused by misconfigured access controls, and added mandatory peer review for production access24 | Scoped roles and peer review for production changes |
| Amazon Q extension, July 2025 | An over-scoped token let an attacker ship a destructive prompt in a public release of the extension; AWS pulled the version and says the code could not execute25 | Least-privilege CI tokens and protected release branches |
| Nx "s1ngularity", August 2025 | Malicious package versions drove locally installed AI command-line tools to hunt for secrets, and more than 2,000 verified secrets leaked26 | Pinned dependencies, install scripts off, no permission-skipping mode on machines with secrets |
| Claude Code, 2025 to 2026 | A malicious repository's settings file could run commands or redirect the user's API key before the trust prompt appeared; both flaws were patched27 | Untrusted repositories opened only in disposable containers |
| IDEsaster, December 2025 | More than 30 vulnerabilities across AI coding tools, with one attack chain that worked against every major AI-integrated IDE tested, leading to data theft or code execution28 | A sandbox with an outbound allowlist, and approval for configuration changes |
| GitHub MCP, May 2025 | A malicious public issue steered an agent holding a broad token into leaking private repository data into a public pull request29 | Fine-grained tokens scoped to one repository |
| Committed screenshots, September 2026 | Researchers found more than 13,000 internal images from developers at over 300 organizations, including billing records, committed to public repositories by coding agents30 | Review of every agent commit, and controls on repository visibility |
The pattern is consistent. The most damaging incidents happened because an agent could reach production or the backups, and cloning a repository has become a code-execution event because the repository itself can carry the payload. No measured study yet shows how much sandboxing or required review lowers incident rates, so the case for these controls rests on the incidents themselves, and the controls are cheap next to one lost database.
Seats became meters
Adoption is settled. Stack Overflow's figures show AI tool use among developers rising from 44% in 2023 to 79% in 2025, and agent use climbing from 31% to 59% by an April 2026 pulse survey, while positive sentiment fell.31 The money has changed shape with it. GitHub moved Copilot to token-based credits in June 2026, explaining that under seats a quick chat question and a multi-hour autonomous coding session could cost the same, and added budget caps at enterprise, cost-center and user level.7
Uber ran the whole cycle in public. It reportedly encouraged staff to use AI as much as possible, used its 2026 AI budget in about four months, then capped monthly spend per employee for each agentic coding tool, with exceptions by permission.8 Meta told engineers in September 2026 that it will not use AI adoption dashboards or token counts to evaluate impact.9 Gartner expects the pressure to grow: it predicts inference cost per agentic workflow will rise more than fivefold through 2028 even as unit prices fall, and recommends routing tasks to the right model tier.32 Our guide to inference cost covers those controls.
The cost drivers are tokens per task from long context, reasoning and retries, the default model tier, session length and autonomy, overlapping tools per engineer, and above all incentives. Budgeting becomes a consumption forecast with a long tail of heavy users, and usage can never be the success measure, because usage is exactly what the meter charges for.
The harness is the unit of control
The credible practitioner sources describe one operating model. Birgitta Böckeler's account of harness engineering defines the agent as the model plus its harness, and divides controls into guides that steer before the agent acts and sensors that check after, some deterministic like tests, linters and type checkers, others inferential like AI review.33 Anthropic's guide for Claude Code makes verification its first practice, giving the agent a check it can run, and draws the line policy depends on: context-file instructions are advisory, while hooks are deterministic and always run.34 GitHub builds separation of duties into its cloud agent, which pushes only to its own branch, opens pull requests that a human must review and merge, does not trigger workflows until review, and prevents the person who asked for the change from approving it.35
Context files such as AGENTS.md, an open format now used by more than 60,000 projects, are useful guidance.36 Governance belongs in permissions, sandboxes, hooks, branch protection and CI, which the agent cannot talk its way past. The table maps the clauses that appear in coding-agent policies to the control that enforces each one.
| Policy clause | Enforced by | Why |
|---|---|---|
| Agents never touch production | Agent identities with no production roles, short-lived environment-scoped tokens, backups out of reach | Replit, PocketOS |
| Agents run contained | An operating-system sandbox, an outbound network allowlist, no permission-skipping mode where secrets live | IDEsaster, Nx |
| Untrusted repositories open safely | Disposable containers, with a trust prompt before any repository configuration runs | Claude Code flaws |
| Someone other than the requester approves | Agent-only branches, draft pull requests, branch protection, CI held until review | GitHub's built-in model, Amazon Q |
| Generated code meets the security bar | Static analysis with data-flow tracking, dependency and secrets scanning as hard gates | Veracode 56% |
| Dependencies are real and vetted | Lockfiles, an internal registry or proxy, new packages blocked by default | Invented package names |
| Tools and MCP servers are approved | An allowlisted registry with pinned versions, and no lethal-trifecta combinations | GitHub MCP exploit |
| Agent work is attributable | Agent-authored, signed commits with a human co-author, and audit logs | Dispute and incident response |
| Changes stay small | Pull request size limits in the review policy | DORA's small-batch finding |
| Spend is bounded | Budgets per user and per tool, pooled credits, routing by task | Uber, Gartner |
Open adoption
Every tool, every setting
- Fast uptake and high enthusiasm
- Agents hold whatever the developer holds
- Usage leaderboards and surprise bills
- Review queues and incidents absorb the gain
Ban or restrict
Approved autocomplete only
- Lowest immediate risk
- Developers move to personal tools anyway
- Forgoes documented gains on migrations and tests
- Leaves the organization behind on skills
Harnessed
What we run
- Agents contained and scoped by default
- Every change reviewed by someone other than the requester
- Hard CI gates on tests, security and dependencies
- Cost and stability reported beside throughput
Measure delivery and cost together, never usage alone
METR's perception gap rules out self-reported time savings, and Meta's reversal rules out token counts. DORA's delivery metrics and DX's framework of utilization, impact and cost point to one scorecard that pairs throughput with stability, adds the review queue where the cost now lands, and sets spend against results.337 Baseline it before the rollout, because a team cannot tell an improvement from a dip it never measured.
| Dimension | Metrics | What it catches |
|---|---|---|
| Throughput | Lead time for changes, deployment frequency, merged pull requests per engineer | Whether more code becomes more releases |
| Stability | Change failure rate, recovery time, rework, split for agent and human changes | The instability DORA measures |
| Review flow | Review latency, pull request size, reviewer load | The queue where the gain is lost |
| Security | New findings per merged change by severity, secrets caught before merge | Insecure generated code |
| Cost | AI spend per team per month, set beside the change in throughput | Spend growing faster than results |
| People | Developer experience, junior debugging skills | Skills that stop forming |
Where to start
The best-documented enterprise wins share one shape: a bounded job, a mechanical check and a success threshold set in advance. Google set a goal of saving at least 50% of end-to-end time for its code-migration tool, and 80% of the landed changes in one migration were AI-authored.38 Airbnb migrated about 3,500 test files in six weeks against a manual estimate of a year and a half, through a pipeline that refactored, tested, linted and compiled with retries.39 Uber's AI code reviewer now analyzes more than 90% of about 65,000 weekly changes, and engineers rate 75% of its comments useful.40 Upgrades of this kind are covered in our guide to AI-assisted .NET and Java modernization.
- Lock down identity and containment Give agents identities with no production roles, short-lived scoped tokens, sandboxes with outbound allowlists, and backups they cannot reach. This needs no productivity evidence to justify.
- Put agent output on existing rails Agent-only branches, draft pull requests, an approver who is not the requester, and hard gates for tests, security, dependencies and licenses. Write size limits into the review policy.
- Choose work you can verify Start with migrations, framework upgrades, test backfills and AI-assisted review, set the success threshold first, and expect a dip in the first months.
- Govern cost as consumption Budgets per user and per tool with an exception path, fewer overlapping tools, routine tasks routed to cheaper models, and token counts kept out of performance reviews.
- Read the scorecard honestly If pull requests climb while deployments and change failure rate stay flat, the fix is in review, testing and release. Keep junior engineers debugging by hand often enough to learn.
Questions to settle before scaling coding agents
- Can any agent, on any machine, reach production data, infrastructure or backups today?
- Which tools and MCP servers are approved, and where is that enforced?
- Who approves an agent's pull request, and can the requester approve their own?
- Which CI gates block a merge, and can an agent change them?
- What is each team's monthly AI budget, and who sees spend beside delivery results?
- Which baseline will show whether delivery, stability and security improved?
The question has moved from whether AI writes code faster to whether the organization can absorb what it writes. Generation is abundant now. Verification, review and safe release are scarce, and the meter runs on every token whether or not the result ships. An engineering organization adopting agents is really buying a verification system, an identity model and a consumption budget, with the model as the most interchangeable part.
This is how we run custom software development. Our engineers work with coding agents inside a harness: agents contained and scoped, every change reviewed by an engineer other than the one who asked for it, hard gates on tests, security and dependencies, and delivery, stability and cost reported together, so clients get the speed and keep a codebase they can trust.
Questions leaders ask
Do AI coding agents make developers more productive?
For individual output on well-scoped, greenfield or repetitive work, yes. Trials range from 26% more completed tasks in large enterprise experiments to a 19% slowdown for experts in their own repositories. Across more than 500,000 developers, commits rose about 240% with agents while releases rose about 30%, so team delivery gains depend on review, testing and release keeping up.
Is AI-generated code secure?
Often not by default. Veracode's 2026 tests found generated code passing security checks 56% of the time on average, with cross-site scripting passing 15% of the time, and newer models barely improved on the year before. Treat every generated change like any other untrusted contribution, with static analysis, dependency and secrets scanning, and human review as merge gates.
What controls should an enterprise require for coding agents?
Agents with no production credentials, sandboxes with outbound allowlists, untrusted repositories opened in disposable containers, agent-only branches with approval by someone other than the requester, CI held until review, hard security and dependency gates, an approved list of tools and MCP servers, signed and attributable commits, and per-user budgets.
Is a CLAUDE.md or AGENTS.md file enough to govern an agent?
No. Context files steer an agent and are useful, and Anthropic's own guidance calls those instructions advisory. Rules that must always hold belong in deterministic controls the agent cannot override: permissions, sandboxes, hooks, branch protection and CI checks.
How should we measure the impact of AI coding tools?
Measure delivery and stability together, using lead time, deployment frequency, change failure rate and recovery time, plus review latency, pull request size, security findings per change and AI spend per team. Avoid self-reported time savings and token counts; METR found developers believed they were faster when they were slower, and Meta dropped token counts from performance evaluation.
Why are AI coding costs rising when token prices fall?
Agents use far more tokens per task through long context, reasoning, retries and long autonomous sessions, and vendors have moved from seats to metered pricing. Gartner predicts inference cost per agentic workflow will rise more than fivefold through 2028. Budget per user and per tool, route routine work to cheaper models and report spend beside delivery results.
Which work should we give coding agents first?
Work with a mechanical check: code migrations, framework and language upgrades, test backfills and AI-assisted code review. Google, Airbnb and Uber all published measured results on jobs of this shape. Open-ended feature work in large legacy systems is where trials show the smallest gains.
Sources
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityMETR, arXiv 2507.09089, July 2025
- Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding ToolsNational Bureau of Economic Research, May 2026, revised September 2026
- Announcing the 2025 DORA ReportGoogle Cloud, September 2025
- 2026 GenAI Code Security: Syntax is Solved, Security is NotVeracode, July 28, 2026
- Vibe coding service Replit deleted production databaseThe Register, July 21, 2025
- 'It took 9 seconds': tech founder outlines how rogue Claude-powered AI tool wiped entire company database and backupsTechRadar, May 2, 2026
- GitHub Copilot is moving to usage-based billingGitHub Blog, 2026
- Uber caps employee AI spending after blowing through budget in 4 monthsTechCrunch, June 2, 2026
- Exclusive: Meta Tells Engineers AI Token Usage Won't Be Part Of Performance ReviewsThe Information, September 2026
- The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software DevelopersMicrosoft Research
- How much does AI impact development speed? An enterprise-based randomized controlled trialarXiv 2410.12944, Google
- We are Changing our Developer Productivity Experiment DesignMETR, February 24, 2026
- Will AI Replace Software Engineers?Yegor Denisov-Blanch, Stanford, conference slides, 2025
- Sundar Pichai shares news from Google Cloud Next 2026Google Blog, April 22, 2026
- Why SWE-bench Verified no longer measures frontier coding capabilitiesOpenAI, February 23, 2026
- AI Impact on Engineering Productivity: 2026 Report DataFaros AI, 2026
- Complacency with AI-generated codeThoughtworks Technology Radar
- AI | 2025 Stack Overflow Developer SurveyStack Overflow
- How AI assistance impacts the formation of coding skillsAnthropic, 2026
- Introducing DORA's inaugural AI Capabilities ModelGoogle Cloud, September 24, 2025
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsarXiv 2108.09293, 2021
- We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMsarXiv 2406.10279, USENIX Security 2025
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationSimon Willison, June 16, 2025
- AI coding bot didn't take down AWS, Amazon confirmsAbout Amazon, 2026
- Security Update for Amazon Q Developer Extension for Visual Studio Code (Version #1.84)Amazon Web Services, July 2025
- s1ngularity's aftermath: analysis of Nx supply chain attackWiz, August 2025
- Check Point Researchers Expose Critical Claude Code FlawsCheck Point, 2026
- Critical flaws found in AI development tools are dubbed an 'IDEsaster'; data theft and remote code execution possibleTom's Hardware, December 2025
- GitHub MCP Exploited: Accessing private repositories via MCPInvariant Labs, May 2025
- AI Coding Agents Exposed 13,000 Internal Images, Including Billing Records, on GitHubThe Hacker News, September 2026
- Getting ready for 2026 results: A look back on Developer Survey findingsStack Overflow Blog, September 30, 2026
- Gartner Predicts AI Inference Costs Per Agentic Workflow Will Increase More Than Fivefold Through 2028Insurance-Canada.ca, September 24, 2026
- Harness engineering for coding agent usersBirgitta Böckeler, martinfowler.com, April 2, 2026
- Best practices for Claude CodeAnthropic, Claude Code documentation
- Risks and mitigations for GitHub Copilot cloud agentGitHub Docs
- AGENTS.mdagents.md
- How to measure AI's impact on developer productivityDX
- Accelerating code migrations with AIGoogle Research
- Accelerating Large-Scale Test Migration with LLMsAirbnb Engineering, March 13, 2025
- uReview: Scalable, Trustworthy GenAI for Code Review at UberUber Engineering, 2025
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 4, 2026. No client data appears in our insights.