In May 2026, the Bun project completed a code migration with almost no precedent in software-development history: 535,000 lines of Zig, 11 days, 64 Claude instances working in parallel, 6,778 commits, 1 million lines of code changed. API cost: $165,000 — roughly one year's output for three engineers.
The figure ignited fierce debate on Hacker News — some ran the numbers and called it a bargain; others said it downplayed the real investment. But they were all asking the wrong question. The real question is not "was $165,000 worth it?" It is: who reviews 1 million lines of code?
The Popular Narrative Misses a Key Fact
Most coverage framed the Bun rewrite as an "AI coding efficiency revolution" — 64 Claudes working simultaneously, 1,300 lines per minute, a year of human work in 11 days. Enticing, but it omits a critical detail. Jarred Sumner candidly disclosed in his blog: roughly 4% of Bun's Rust code sits inside unsafe blocks — about 13,000 unsafe keywords across ~27,000 lines. For comparison, uv at ~350,000 lines has only 73 unsafe calls. Bun's count is 178× higher. And undefined behavior surfaced in "safe" Rust too — which is harder to debug than C++, precisely because you believe safe code cannot break.
Sumner also admitted the rewrite introduced 19 known regressions, mostly from syntactically identical but semantically different code — Zig's assert is a function whose arguments run on every build; Rust's debug_assert! is a macro whose expression is deleted in release builds. Two code blocks that look identical behave completely differently. The regressions were fixed. But "fixed" does not mean "no other problems exist."
A million lines of change cannot be human-reviewed line by line. At one line per minute, that is 11.7 days of continuous reading; at a realistic review speed of 200 lines per hour, it takes over two years. The PR's actual reviewers were primarily claude[bot] and coderabbitai[bot]. Sumner's own review approach: "check that the adversarial review agents correctly caught divergences, ensure conversion guidelines were followed, and personally read a fair amount of code." How much is "a fair amount"? He did not say.
This is the most noteworthy signal in the entire story: AI-written code will ultimately be reviewed only by AI.
Why Bun Had to Rewrite — and Why It Couldn't Before
Bun's original code was Zig, which — like C — does not automatically manage memory. JavaScriptCore, the JS engine Bun embeds, has extremely strict GC and exception-handling rules. When those two paradigms coexist in one process, every memory allocation requires line-by-line audit: where are these bytes freed? How do you ensure single-release? Is this GC memory or manually managed?
The team tried everything: Address Sanitizer support on the Zig compiler, ASAN tests in every CI run, ReleaseSafe builds on Windows, 24/7 Fuzzilli fuzzing, extensive end-to-end memory-leak tests. Crash reports still poured in. "Our bug-fix backlog feels awful; I'm tired of going to sleep worrying about Bun crashes," Sumner wrote. The Rust version's answer: 2,000 iterations of Bun.build() dropped memory from 6.7 GB to 609 MB; 128 reproducible bugs fixed; binary size down ~20%.
This was never a "should we rewrite" question — the fundamental conflict between GC and manual memory management left no choice. But why couldn't it happen before? Because rewriting meant freezing a year of bug fixes, security patches, and new features. An open-source project with zero revenue could not absorb that cost. Claude Fable 5 changed the equation — not by "being able to write code," but by enabling Sumner to design an entirely new working method.
Implementer/Reviewer Separation: the AI-Native Development Methodology
This is the most underrated part of the Bun rewrite — Sumner's methodology, not the AI's capability. He decomposed the process into roughly 50 dynamic workflows, each a loop: one Claude writes code based on context (Jira ticket or GitHub issue); two reviewer Claudes review it; feedback is applied; the next task is claimed. At peak, Sumner ran 4 concurrent workflows × 16 Claudes each = 64 instances working across 4 work trees simultaneously.
The critical design: implementer and reviewer are completely separated. The writing Claude wants its code accepted — the same bias any human engineer has. So reviewers see only the diff, never the implementer's reasoning, and are explicitly instructed: "assume the code is wrong." Each implementer faces two or more adversarial reviewers whose sole job is finding bugs. This is not "having AI write code." This is designing a software-development organization composed of AI — implementers, reviewers, integrators, each with defined roles and constraints.
Zig code was a single compilation unit; Rust splits into ~100 crates. The first cargo check produced ~16,000 errors. For a human, catastrophic. For 64 parallel Claudes, a manageable work queue: the workflow grouped errors by crate, ran cargo check per crate, one Claude fixed, two reviewed, one applied. Two days later, Linux's failing tests dropped from 972 to 23. A day and a half after that, Linux was green. Five days in, all six platforms passed. The test suite skipped and deleted nothing.
The Transfer of Code-Review Rights
The Bun rewrite reveals a structural shift: the transfer of code-review rights. Traditional development treats code review as the last line of quality defense — a PR lands, at least one human engineer reads every line, confirming logic, security, and style. That mechanism assumes humans can review all code and have time to do so. The Bun rewrite broke both assumptions simultaneously. A million lines cannot be human-reviewed. And the deeper question: even with time, can humans meaningfully review AI-generated code? AI code follows correct syntax and type constraints but may do semantically unexpected things. A human reviewer facing "looks correct" code loses the intuitive "something smells wrong" sense that human-written code triggers. That is why reviewers must also be AI — explicitly instructed to assume the code is wrong.
The transfer runs in three stages. Stage 1: AI-assisted human review (current mainstream) — AI does initial checks, humans make final decisions. GitHub Copilot and CodeRabbit live here. Stage 2: AI reviews AI-generated code (Bun's current state) — AI generates, AI reviews, humans review the AI's review. The PR's primary reviewers are bots; Sumner's role is "checking the reviewers' work." Stage 3: AI generates, reviews, and maintains — code from generation through maintenance entirely by AI, humans intervening at key decisions only. Post-acquisition, Bun partially operates here already: the only tool that can effectively maintain this codebase is Claude itself.
What the Frame Explains
The "code-review rights transfer" frame illuminates several phenomena. Open source's maintainability assumption is being upended. Open source's core premise: code is public, therefore anyone can review, fork, and contribute. But when a codebase's scale reaches a point only AI can understand, "public" does not equal "maintainable." Community members already note Bun "is no longer a traditional open-source project — to submit a PR, you need an Anthropic subscription." Ownership of "AI-native code" is unsettled. If code is generated, reviewed, and maintained by AI, who is the project's "owner"? The prompt writer? The compute provider? The party with the strongest model? Bun's acquisition by Anthropic is no coincidence — when your codebase can only be maintained by Claude, you are bound to Anthropic. The frame maps across domains. AI-generated legal documents, financial reports, medical diagnoses — when AI content reaches a volume humans cannot review item by item, who guarantees quality? The answer, as with Bun, may be: another AI. But the AI-reviews-AI loop raises a fundamental trust problem: what if two AIs collude? What if they share the same training-data biases?
What It Leaves Us
The Bun rewrite is not a case of "AI replacing programmers." It is a warning: a widening chasm separates efficiency gains from maintainability. If you are a technical decision-maker, look past API cost — the $165,000 is the entry ticket. The real cost is "maintainability debt": when your codebase can only be maintained by a specific AI, your technology-stack choice becomes a model choice. Choosing Claude today and switching to Gemini tomorrow may cost more than Zig-to-Rust did. If you are an open-source maintainer, think hard about AI-generated code review strategy — do not assume "CI passed, so it's safe." Bun's tests all passed, yet 19 regressions and 13,000 unsafe blocks remained. Consider adversarial review workflows — not AI writing code, but AI reviewing code. If you are a developer, stop focusing on "writing better prompts." The Bun rewrite's most valuable part is not the AI's capability but Sumner's workflow architecture: implementer/reviewer separation, adversarial review, parallel work trees. Learning to design AI collaboration systems is 100× more important than learning to write prompts.
The Bun rewrite's ultimate legacy is not the $165,000 bill or the 6,778 commits. It is a question: when code no longer requires human understanding, do we still need human developers? The answer may be no — but only if we accept a software world maintained by AI.
References: Jarred Sumner, "Rewriting Bun in Rust" (bun.com/blog/bun-in-rust) · InfoQ/36Kr coverage, July 2026.
