For years, the hidden reasoning chains behind proprietary frontier models have been treated as the most closely guarded asset in the AI industry. Vendors encrypt them, bill for them separately, and expose only the polished final answer. A new research paper now demonstrates that this lockbox has a structural flaw — and that a cheap, lightweight model can be repurposed as the key.
The anti-distillation moat, long considered the core defensible advantage of closed frontier models, does not hold up in practice.
What Happened: An End-to-End Bypass, Not a Theory
In August 2026, a paper titled Stealing Reasoning Traces from Proprietary LLM APIs sent ripples through the research community. The team behind it spans the Max Planck Institute for Intelligent Systems, the ELLIS Institute in Tübingen, and the security firm Snyk. Their finding is blunt: the encrypted chains of thought returned by top-tier models from Anthropic, OpenAI, and Google can be recovered in plaintext — with the help of the vendors’ own smaller models.
This was not a speculative attack sketch. The researchers walked the full path in a live demonstration: encrypted reasoning blocks captured from Claude 3 Opus 4.8 were replayed into Claude Haiku 4.5, and a jailbreak prompt coaxed the smaller model into transcribing the flagship model’s raw reasoning, word for word.
The validation method is what makes the paper hard to dismiss. The team used API billing as a measuring stick: the number of plaintext tokens extracted from the decrypted traces closely matches the number of hidden reasoning tokens each API call was charged for, hugging the y=x line on their reconciliation plots. In other words, Haiku really was emitting Opus’s thinking process, not improvising a plausible imitation.
Anatomy of the Flaw: Three Layers
Layer 1: Cross-Model Portability
The encrypted trace format was designed for interoperability inside each vendor’s ecosystem — and that is exactly its weakness. The same ciphertext can be replayed across models, sessions, and users within one provider. Because serving infrastructure is largely stateless, vendors often reach for globally shared encryption keys to keep storage costs down. With the exception of Anthropic’s Fable 5, the researchers found that the Claude family, the GPT-5.6 series, and the entire Gemini line all allowed low-tier models to read the encrypted chains of thought of their high-tier siblings.
Crucially, this is not a break of the underlying cryptography. It is an abuse of data-format interoperability between a vendor’s own models, using the weakest member of the family as the point of entry.
Layer 2: The Economics Are Absurd
At Haiku 4.5 API pricing, the paper estimates the cost of decoding ten thousand complete reasoning chains at roughly 720 US dollars. A one-time, low-four-figure budget is enough to bulk-extract the core reasoning logic of frontier models at industrial scale. A defense that can be overwhelmed with a modest credit card balance is not much of a defense.
Layer 3: Data Fingerprints and the Distillation Question
The most contentious part of the paper is its memorization analysis. The team probed Kimi-K3, GLM-5.2, DeepSeek-V4-Flash, Kimi-K2.6, and Inkling with prompts derived from the stolen traces, and the results are striking:
- A verbatim-reproduction cost gap. The theoretical query cost for Kimi-K3 to continue an Opus reasoning trace for its next 16 tokens sits 4 to 6 orders of magnitude below DeepSeek-V4-Flash and Inkling.
- A 1% prefix bends the style. Seeding just the first 1% of an Opus trace as context measurably pulls Kimi-K3’s final output toward Claude’s stylistic fingerprint.
- A swapped-prefix control stays flat. When Kimi-K3 was fed Inkling’s reasoning prefix instead, nothing happened — the curves only align where the Opus lineage is involved.
The authors are careful to note that these experiments demonstrate anomalous behavioral compatibility, not provenance of training data. But the asymmetry in the numbers is exactly the signature the industry’s anti-distillation defenses are supposed to prevent — and its exposure shows how fragile those defenses are.
From ‘Encryption Equals Safety’ to ‘Transparency by Design’
The deeper lesson challenges an industry default: that hiding the reasoning process is the moat. The traditional logic goes like this — seal the model’s intermediate thinking, ship only the final answer, and the core cognitive capability stays black-boxed. The paper shows that this assumption fails at the engineering layer, where convenience and cost control quietly override cryptographic hygiene.
| Old assumption | New reality |
| Hidden chain of thought = core competitive advantage | Hidden chain of thought = an extractable asset |
| Anti-distillation = a technical barrier | Anti-distillation = a barrier with a documented bypass |
| Closed source = safe | Closed ≠ safe; transparency discipline matters more than secrecy |
The shift is conceptual as much as technical: an asset you cannot isolate is not a moat, it is a liability with extra steps.
What Vendors Should Do Now
The fixes follow directly from the three layers of the flaw. Kill global master keys and move to per-session, per-user encryption for reasoning traces. Make low-tier models cryptographically incapable of decrypting high-tier traces, so a jailbreak on Haiku yields nothing from Opus. Treat replay-shaped traffic — the same ciphertext block appearing across sessions — as a first-class anomaly signal, not an accounting curiosity. And disclose the threat model to enterprise customers who are currently billed for hidden tokens under the assumption that they are private.
None of this requires new science. It requires accepting that the adversary is no longer only outside the API boundary — it is the cheapest model in your own catalog.
The Takeaway
The moat is shifting. It used to be: hide the reasoning, sell the answer. After this paper, hiding is no longer a strategy — making reasoning trustworthy, auditable, and priced honestly is. The vendors affected now face a choice between quiet patching and a transparent account of what was exposed. For everyone building on these APIs, the next few weeks of vendor responses will be worth reading as carefully as the paper itself.