Chain-of-Thought Extraction Breaks AI Reasoning Moats

Imagine paying a few hundred dollars to read the private thoughts of a model that cost billions to build. That is now possible. A 116-page study from MATS, ELLIS Tübingen and collaborators demonstrates that the internal reasoning of Claude, GPT and Gemini can be pulled out in plain text using two API calls — performed by each company’s own cheapest assistant model. At about $720 per ten thousand traces, the price tag of a once-impossible extraction is now trivial. And that triviality is precisely what makes this a strategic story, not just a security bulletin.

The weakness lives in design, not in the cipher

Focusing on “the encryption failed” misses the point. The real story is a storage trade-off. Reasoning tasks need continuity — one step must inherit the context of the last. Holding that context on the server is costly and collides with no-retention promises, so vendors hand the encrypted reasoning to the client for safekeeping, retrieving it on the next request. The shortcut is where the exposure was born.

The blocks have no binding to a session, a user, or a generating model. They can be moved around freely inside the vendor’s ecosystem. Interoperability was the goal — smoother downgrades, compressed history, graceful switching. But it also gave the cheapest, least-guarded sibling the ability to recite the flagship’s thinking. When a defense ignores its own weakest point of entry, it was never a defense at all.

A re-pricing of “hidden intelligence”

The commercial reading matters more than the exploit. If anyone can bulk-extract hidden reasoning for a few hundred dollars, the premium attached to “invisible thinking” loses its basis. The same paper also converts the distillation rumor into a measurable phenomenon: one model reproduced another’s phrasing at up to a million times the expected rate, and similarity scores jumped from 0.17 to 0.33 in priming tests. Copycat training stops being a rumor the moment it can be measured — and measurable copying carries legal risk.

Three audiences, three takeaways

For vendors, the patch is binding, replay-proofing, or moving state back to the server with short-lived storage. The bigger question is what remains after “secret thinking” is gone — for most labs, the answer is data, tooling and ecosystem, the parts a cheap API call cannot reach.

For developers, treat agent logs as credentials. Reconstructing 315,000 reasoning blocks from 6,708 public traces surfaced 328 blocks — roughly 4.9% — that leaked keys, passwords or tokens. Rotate credentials, prefer short-lived ones, and scrub logs.

For model buyers, the hidden-reasoning gap between closed and open weights is smaller than benchmark tables suggest. Choose on data, tooling and alignment fit rather than paying a premium for opacity.

The companion piece Chain-of-Thought Extraction Attack on Claude & GPT details the exploit mechanics and the defensive checklist; this article covers the business and industry impact. Read them together for the full picture.

Related reading: DeepSeek V4 Pro API guide, how MiniMax co-designs models and harnesses, and our analysis of Qwen 3.8 open weights; the paper is arXiv 2608.09867.

Related News