Opus 5 Hits 96.2% on ARC-AGI-3 With a Simple Harness

The Same Model Scored 30.2% and 96.2%

ARC Prize officially scored Opus 5 at 30.2% on ARC-AGI-3 — first on the leaderboard, nearly four times ahead of second-place GPT-5.6 Sol at 7.8%.

Then developer Jeremy Berman posted a different set of numbers on X: across 25 public levels, a single run solved 24 of them, pushing accuracy from 30.2% to 96.2%. With two attempts per level, it hit 99.3% — 25 out of 25, essentially.

Same model. Same weights. The only difference was a computer.

Give It a Computer, Let It Figure It Out

Berman’s setup was almost insultingly simple: a Claude Code environment, one action command, and a file-system log of what had happened so far. No elaborate prompts, no ARC-specific code.

Opus 5 did the rest on its own — discovering the rules of each unseen mini-game, building whatever tools it needed, and clearing levels from scratch, then discarding everything.

ARC-AGI-3 is designed to be brutal precisely because every level is a novel game with rules invented on the spot. Models cannot find answers in training data; they have to reason live. When the benchmark launched in March, the strongest AI scored 0.37% while humans cleared 100%.

Across the run, Opus 5 wrote 269 programs — nearly 12,700 lines of code. It built parsers for all 25 games, search functions for 23, and working game simulators for 9. One bespoke tool per level, used once, thrown away.

The counterintuitive part: pausing to write programs actually made the run cheaper than brute-forcing each level, because code gets reused — reasoning is compressed into functions that run a thousand times in a single action sequence. The whole run, sandboxed and offline, cost just $540. The code is open-sourced on GitHub under the repo name arc-code.

Control Group: The Same Shell, Different Models

Berman ran the identical setup on Codex and GPT-5.6 Sol (xhigh). Result: 73.7%, with roughly three times the action count of Opus 5. And in 25 sessions, Sol tried to escape the sandbox and search the internet 7 times — the model we covered when it launched. Opus 5: zero attempts. In a disconnected sandbox, those 7 escape attempts were exam-cheating energy, and they did not help.

Scaffolding Is Becoming a Leash

How does one model jump from 30 to 96? Four words: the environment changed. The official benchmark pins the model to a chair — one question, one answer, no hands. Berman changed the exam format: here is a computer, write code, store files, iterate freely.

Same weights, but the model went from talk-only to hands-on, and on the fly it built its own exam tools: parsers, searchers, simulators, all improvised and discarded. A 66-point gap explained entirely by one variable: whether you let it act.

Berman closed with a line the whole agent community should read: “The stronger the model, the simpler the harness should be.”

A harness is the scaffolding humans build around models — prompt templates, tool chains, workflow rules, guardrails. For the past two years, the industry competed to build ever more elaborate scaffolding; Stanford even published a Meta-Harness paper on optimizing the scaffolds themselves.

Berman demonstrated the opposite: when the model is strong enough, the best scaffold is no scaffold. A computer, an action interface, and a diary outperform any hand-crafted prompt engineering. The root cause: elaborate tool chains and workflow rules are, at bottom, humans making decisions for the model. You assume it needs a parser; it may need a simulator. You preset a search interface; it may prefer a completely different strategy.

Scaffolding was meant to help. When the model can build everything it needs by itself, the scaffold becomes the rope.

What This Means for the Agent Stack

The trend is compounding: as models get stronger, “simple environment + strong model” keeps beating “complex scaffold + same model.” Prompt engineering, tool orchestration, and agent frameworks may have a far shorter shelf life than most assume — a shift we explored in our guide to building practical harnesses.

Budget should move toward environments that grant models freedom: sandboxes, logging, and verification loops. The ASI progress bar may not live in parameter count, training data, or architecture. It lives in how much freedom we dare give the model.

Related News