Sparse Weight Decomposition: LLM Circuits at 1% Data Cost

Interpretability has carried an embarrassing contradiction for years: to understand a model that was already trained, researchers had to train another model. Methods like Transcoder and sparse feature modules learn a fresh set of internal units to approximate one layer, then hunt for task circuits by keeping or removing those units. To open a black box, you first had to build a second, slightly smaller black box.

A collaboration between IQuest Research, the Safe AI Forum, Oxford, Stanford and Tsinghua takes a more direct route. Their method, Sparse Weight Decomposition (SWD), skips the replacement model entirely. It factors a pretrained dense weight matrix into two sparse matrices, and turns the shared middle dimension into bottleneck units that can be scored, selected and ablated directly on the weights. The result is circuit extraction that uses under 1% of the data that trained baselines need, verified from GPT-2 all the way up to Qwen3.5-27B.

The hidden cost of opening a black box

The old approach has a real price, and not just in compute. To approximate a module you need data and optimization, and if your substitute module does not faithfully reproduce the original, every subsequent analysis is really explaining both the model and the replacement error at once. Task circuits also demand units that can be intervened on independently: keep only them and the capability survives; remove them and performance collapses. Training-based substitutes usually have to be fitted separately per model, per layer and per checkpoint, which makes large-scale, systematic circuit analysis painfully expensive.

The paper frames the ideal as three requirements at once: preserve the original model’s behavior, provide independently intervenable units, and build them cheaply from weights that already exist. SWD is designed to hit all three.

Decomposing weights into sparse read-write paths

For a dense weight matrix W, SWD finds two sparse matrices A and B such that W ≈ A·B. A and B share m intermediate dimensions: the i-th column of A reads a scalar out of the input, and the i-th row of B writes it to the output. One column plus one row forms a bottleneck unit — a fixed rank-one read-write path.

Think of it as adding transfer stations to an extremely dense road network: each station only connects a handful of entrances and exits. You can see where a path reads from and writes to, and you can switch a single path off to watch what the model does. During decomposition, SWD uses a small amount of calibration text to minimize the output error before and after the split, alternates between updating the two factors, applies hard thresholds to hold a nonzero budget, and refits the surviving values.

Three results stand out:

  • Less than 1% of the data. On a GPT-2 Small layer-8 matrix, SWD enters a low cross-entropy-error regime with just a few thousand calibration tokens, where Transcoder and VPD-Recon-CI need on the order of 10⁶ optimizer-replay tokens to get close. The same pattern holds on Qwen2.5-0.5B/1.5B/3B and Qwen3.5-27B.
  • Fewer units for the same causal test. On GreaterThan, IOI, Docstring and Gendered Pronoun tasks, SWD reaches the same sufficiency and necessity targets with fewer bottleneck units and active connections than trained baselines.
  • Full-model and zero-data versions. SWD replaces all 48 attention and MLP matrices of GPT-2 Small; the fine-tuned variant (SWD-FT, keeping sparse structure and refitting nonzero values) drops model CE from 3.90 to 3.44 — using ~20.6 million tokens versus 2.884 billion for a sparse-pretraining baseline at similar CE, again under 1%. A zero-data variant needs no calibration text at all, minimizing only the Frobenius error between original and decomposed weights, and still yields valid task circuits.

Why not just use SVD?

Singular value decomposition also expresses a matrix as a sum of rank-one components and needs no training data. Recent work (NaNA) even shows SVD components can rank tasks and support intervention. So why does SWD exist?

Because for circuit analysis, few units is not the same as few computational paths. SVD components have dense read and write directions — even a handful of components still touch almost every input and output dimension. SWD pushes sparsity down to each unit’s read-write connections, so a small number of units really does mean a small number of active connections. The paper’s control experiments make the point cleanly: full-rank SVD and a random-orthonormal-basis decomposition both reproduce the original matrix exactly, yet need many more active connections than SWD to reach the same task performance. The win is not “splitting the matrix” — it is the sparse read-write structure of each unit.

Circuit extraction is becoming an audit surface

Mechanistic interpretability has long been treated as a boutique, expensive exercise. SWD attacks the cost curve directly: when you can carve circuits out of any checkpoint at under 1% of the data cost, circuit extraction stops being a research hobby and starts becoming an engineering and audit surface.

Three implications matter:

  • Auditing what a model can do. If you can cheaply trace which circuits drive a capability, you can check whether a model hides abilities that are not advertised. This pairs naturally with the ongoing work on extracting and probing hidden reasoning chains — both are ways of reading what is really inside a model rather than trusting its surface behavior.
  • Precise model surgery. The team shows a single bottleneck unit (c205) can be edited: applying its rank-one read-direction update to raw GPT-2 weights raises the correct-answer margin on “The opposite of up is” by 0.216 while producing side effects (mean last-token KL of 4.02×10⁻⁵) below both a random unit and a rank-4 LoRA baseline. Targeted, low-collateral editing is no longer hypothetical.
  • Open-source weights become an audit asset. When anyone can decompose a released checkpoint and inspect its wiring, openness gains a concrete safety argument — the same logic driving the open-source security turn in frontier labs.

What to do with this

For researchers, the paper (arXiv:2608.03913) comes with open code, released weights and an interactive blog — the fastest way to see SWD is to run the GreaterThan demo yourself.

For teams building on open models: add a weight-side decomposition pass to your pre-deployment audit. It is cheap enough now to be part of the pipeline, not a one-off research project.

For safety and governance folks: watch the zero-data variant. Being able to trace circuits checkpoint by checkpoint and training stage by training stage is the kind of capability that turns “interpretability” from a research question into a monitoring tool.

Related News