記事一覧

Rebuilding the SDLC for AI Throughput

2026年5月14日

Rebuilding the SDLC for AI Throughput

On Heavybit's High Leverage podcast last week, Simon Willison said he's stopped reviewing every line of code Claude Code produces — even for production systems. The agent does straightforward implementation correctly, consistently, and in his style. So he treats it the way he would another team's service: as a semi-black box he trusts until it breaks. That matters because Willison isn't reckless. He has 25 years in the industry, Django on his résumé, and a public track record of advocating for rigorous engineering practice in AI-assisted development. If the most thoughtful public advocate for that practice is drifting toward unreviewed agent output, that's a signal about where the broader industry already is.

The more important thing Willison said on the podcast, though, wasn't about code review. It was about throughput: "If you can go from producing 200 lines of code a day to 2,000 lines of code a day, what else breaks?" The answer is that almost every adjacent process breaks, because almost every adjacent process was built on the assumption that the 200-line baseline was approximately correct.

The Numbers Behind the Shift

A quick look at my own git history from one client project makes the personal version concrete. Over the last 30 days I averaged 18,000 lines of code per day — 484,654 insertions and 62,713 deletions, one engineer driving a handful of agents. That's not 10x the historical baseline. It's closer to 90x. (Yep, the '479,995' is real despite looking like fakery.)

The institutional numbers tell the same story. Stripe's internal coding agents, called Minions, ship roughly 1,300 pull requests per week in production, built on a documented six-layer architecture that draws an explicit line between submission authority and merge authority. Shopify CTO Mikhail Parakhin went on the Latent Space podcast last month and reported that PR merges at Shopify are growing 30 percent month over month, up from 10 percent the prior year, with complexity rising at the same time. His description of the git/PR workflow itself was direct: "clearly the main issue and bottleneck for us."

The cost of that throughput shows up in systemic measurements as well. The DORA 2025 report finds that organizations with 90 percent AI adoption see a 154 percent increase in PR size, a 91 percent increase in review time, and a 9 percent higher bug rate. CircleCI's 2026 State of Software Delivery report shows feature-branch throughput up 59 percent year over year but main-branch throughput down 7 percent. Build success rates sit at a five-year low of 70.8 percent. The widening gap between what's flying into feature branches and what's actually shipping to main has acquired a name in the 2026 discourse: verification debt — the backlog of unreviewed, unvalidated AI-generated code that piles up because humans can't read code at AI speed.

Even the most sophisticated agent harness on the market doesn't catch everything. Anthropic's own April 23 postmortem disclosed three production regressions in six weeks. The worst — a caching bug — slipped past, in their words, "multiple human and automated code reviews, as well as unit tests, end-to-end tests, automated verification, and dogfooding." The arXiv AgenticFlict study from April 2026 measured a 27.67 percent merge-conflict rate across 107,000 AI-generated pull requests, with conflict severity scaling with PR size. Per-agent rates ranged from 15.24 percent for Copilot to 31.85 percent for Codex, with Claude Code in between at 25.93 percent. The variance is on the harness, not the model, which suggests integration discipline is mostly a function of how the agent is wired into the development process rather than which model underlies it.

Where the Bottleneck Moves

Your software development lifecycle — design reviews, sprint planning, code review, QA, deployment, incident response — was built on the assumption that code is expensive to produce. Every ceremony, every approval gate exists because a mistake at the coding step burned weeks of engineer time. That assumption no longer holds, and what most organizations have done about it is essentially nothing. They've handed engineers Copilot or Claude Code, watched output velocity jump roughly 5x, and changed exactly zero upstream or downstream processes. The result is a firehose running through plumbing that was sized for a garden hose.

The real question isn't whether agents can write code — they obviously can — but whether the organization around them can absorb what they produce. That question has both an upstream side and a downstream side.

Upstream, Jenny Wen, Anthropic's design leader, has been making the point that design processes were originally built to prevent expensive mistakes. Three months of engineering burned on a bad spec is catastrophic. If building the wrong thing now takes a day instead of a quarter, the risk calculus shifts considerably. You can afford to be wrong; prototype-and-validate beats spec-and-pray when iteration is genuinely cheap. Most engineering organizations have not absorbed this yet, and they continue to run six-week design phases for features an agent team could prototype in an afternoon. The process was built to prevent expensive mistakes, but the mistakes are not expensive anymore. The process is.

Downstream, the traditional signals of code quality have lost much of their information value. When every GitHub repository can have a hundred commits, polished documentation, and comprehensive tests — all generated in 30 minutes — the visible craftsmanship of a repository no longer tells you whether the underlying system has ever survived contact with real users. The same shift plays out inside organizations: when agents produce volumes of well-structured, well-tested code in hours, the bottleneck moves from production to validation. The hard questions become operational. Can you verify the thing works? Can you deploy it safely? Can you operate it at 3 a.m. when it breaks? Does your team understand what the agent built well enough to debug it under pressure?

These are organizational design problems, not tooling problems, and they require engineering leaders who understand both the new capabilities and the process redesign needed to apply them without simply trading one form of technical debt for another.

Practical Moves

A few practical changes follow directly from this picture.

The first is to audit the end-to-end SDLC for throughput mismatch. Walking each step — sprint planning, design review, code review, QA, deployment — and measuring actual cycle time will surface the steps that were designed for a 200-line-per-day baseline and have not since changed. Those are the bottlenecks, and they show up as places where work piles up waiting for a human gate that no longer earns its keep.

The second is to reconsider approval gates. Many were instituted to prevent expensive rework, but if an agent can rebuild a feature in hours, the cost of the gate exceeds the cost of the rework it prevents. Stripe's separation between submission authority (agents can submit) and merge authority (humans hold) is a useful organizing pattern. It concentrates human attention at the one gate that is genuinely irreversible while letting earlier steps run at machine speed.

The third is to shrink the unit of change. Cursor 3.3, released in May 2026, made stacked pull requests the default shape for AI workflows, and stacked PRs are now an industry expectation rather than a leading-edge practice. A 10,000-line mega-PR is a multi-hour reconciliation problem. A stack of fifteen to twenty thin PRs, each under 500 lines and scoped to a single layer (schema, model, API, UI), is reviewable in minutes per layer and can be partially merged as layers stabilize. The AgenticFlict data is consistent with the underlying mechanics: conflict severity scales with PR size, so reducing PR size proportionally reduces the cost of resolving the conflicts that do occur.

The fourth is to shift investment from production to validation. The testing strategy, observability stack, and deployment pipeline are where the new advantage lives. A team that ships and validates ten agent-built features per week will outperform a team that hand-codes two, but only if validation actually keeps pace with shipping.

The fifth concerns code review itself. Reviewing 2,000 lines a day with the same rigor previously applied to 200 is not realistic, and continuing to claim otherwise is mostly performative. The honest move is to redefine what code review is for: architecture review, behavior validation, and integration testing, with semantic-merge tools such as Mergiraf handling the trivial AST-shape conflicts that previously occupied human eyes. Senior engineers should focus on coherence — whether the new code fits the existing architecture, respects existing patterns, and matches the bounded context it lives inside. Coherence is the one thing humans still review better than machines, and it is the one thing Anthropic, Burak Dede, and Bryan Finster have been calling unsolved by the current generation of tooling.

The sixth concerns the design phase. If building is cheap, prototyping is often cheaper than specifying. Building the thing, looking at it, and deciding whether it is the right thing is frequently a better path than exhaustively specifying it in advance. Long PRDs written before the first line of code are a relic of the era when code was expensive.

What's Really Changing

None of this implies that AI coding tools are themselves what separates the winners from the losers in 2026. Everyone has those tools. What separates organizations is which ones have redesigned the processes around the coding step to match the new throughput, and which are still running 2026 velocity through 2022 plumbing.

Code is no longer the bottleneck. The bottleneck has moved to review, verification, and coherence, and rebuilding the processes that handle those is the actual work of AI-native engineering practice.