Claude and Mathematicians Build Order-668 Hadamard Matrix

The smallest open case of a 130-year-old conjecture just fell

For the first time in more than two decades, the list of unsolved Hadamard matrix orders has a new frontier. A team of three mathematicians — Levent Alpöge, Philippe Voinov, and Saul Reynolds-Haertle — working with Anthropic's Claude, constructed a genuine Hadamard matrix of order 668, the smallest order for which no example had ever been found. Epoch AI has provisionally marked the corresponding FrontierMath open problem as "solved by AI." The team did not stop at one: their construction covers all 12 previously unknown orders below 2,000.

What a Hadamard matrix is, and why 668 was the wall

A Hadamard matrix is a square grid of +1 and −1 entries in which every pair of rows is orthogonal — multiply corresponding entries, sum them, and you get zero. Beyond trivial sizes, the order must be a multiple of 4, and the 1893 Hadamard conjecture states that a matrix exists for every multiple of 4. The conjecture remains unproven in general, so mathematicians chip away one order at a time. In 2004, Kharaghani and Tayfeh-Rezaie filled order 428; 668 then became the smallest open case — and stayed open for more than 20 years.

The stubbornness is number-theoretic. 668 = 4 × 167, and 167 is a prime congruent to 3 mod 4 — exactly the class that resists the classical Paley construction. Last year a research team produced a "mod-64" version that satisfied the constraints modulo 64: tantalizingly close to a real matrix, but not one.

How it was delivered: a 23,828-character puzzle

Alpöge announced the result in his trademark cryptic style: a tweet containing 23,828 "+" and "−" characters, with an obfuscated shell script hidden in the replies as a decoder. Decoded, it yields 12 matrices — for orders 668, 716, 892, 1132, 1244, 1388, 1436, 1676, 1772, 1916, 1948, and 1964, precisely the 12 missing orders below 2,000. The independent math database VibeMathed reproduced the decode and re-verified every matrix with exact integer arithmetic: all entries are ±1, and every pair of rows has zero inner product.

The credits read like a 2026 artifact: Alpöge, Voinov, Reynolds-Haertle, and Claude. Alpöge joked that he only claims credit for the bad suggestions. And in a delicious twist, when Menlo Ventures partner Deedy Das asked Anthropic's newer model Fable to decode the tweet, Fable refused — the humans ended up using OpenAI's GPT-5.6 Sol to read the puzzle. Model loyalty, it turns out, is not guaranteed.

Why this is a milestone, not a curiosity

This is the fourth of FrontierMath's 50 open problems to be solved with AI assistance — two by GPT-family models, two by Claude-family models. It follows a remarkable month: the long-studied Jacobian conjecture was overturned by humans working with Fable 5, and an unreleased Claude research model pushed the known lower bound on the Riemann hypothesis significantly higher. A consistent division of labor is emerging: the human mathematician sets strategy, designs verification, and interprets; the model performs tireless combinatorial search in spaces that simply exceed human working memory. It is the mirror image of how researchers recently showed that even cheap models can extract hidden reasoning from frontier systems — the reasoning black box is being opened from both directions, as our chain-of-thought extraction coverage laid out.

What it means

Three things are worth watching.

First, FrontierMath's open-problem set is becoming a public scoreboard for research-grade AI capability — and the pace is accelerating.

Second, the mathematician's job description is shifting. If construction becomes semi-automatable, the scarce skills become orchestration, verification design, and interpretation — exactly what Alpöge demonstrated.

Third, Epoch AI's caveat is the most important sentence in this story: it is not yet clear whether the result comes from an improved search strategy or a generalizable construction. That distinction decides whether this is "one more example" or a new theorem. The technical report will tell.

What to do next

Watch for the technical report, and try reproducing the approach with open-weight models such as the newly released Qwen 3.8 — the search-plus-verification pattern is highly portable. Also note the format: a puzzle-encoded result with independent public verification may become a fast lane for AI-collaborative discovery. And if you benchmark models, "solved a FrontierMath open problem" is now a verifiable capability claim, not marketing.

Leave a Comment

Scroll to top