← Alex's mech-interp place

A deeper analysis of Goodfire's Block Sparse Featurizers

by Alex Jerpelea · August 19, 2026 · code

Abstract

This summer I worked with Amith Ananthram (PhD student in Prof. Kathleen McKeown's NLP lab, currently an Anthropic Fellow) on diffing CLIP against DINOv3, to test whether CLIP's joint language objective costs it fine-grained visual features. Diffing requires good featurizers, and an architecture that really stood out to us is the newly introduced block-sparse featurizer (BSF; Fel et al.). The BSF is similar to an SAE, but its atomic unit is a small subspace (a block of directions) rather than a single direction. It is designed for features that live on low-dimensional manifolds, which are especially frequent in vision.

This work studies the BSF's strengths and weaknesses, finding how it still somewhat suffers from classic SAE failure modes, like feature splitting and composition. We propose several architectural changes to the BSF, and also extend the block paradigm to the crosscoder. We end with a discussion of feature geometry and how BSFs interact with categorical and hierarchical features.

Introduction

In our broader study, we compare the DINOv3 and CLIP models trained in similar settings. DINOv3 is self-supervised just on images, while CLIP is language-supervised, trained to match images to their caption. We expect that CLIP's language alignment makes it lose some vision capabilities. The loose intuition is that humans had acquired vision capabilities long before language during the evolution process.

A big part of diffing models is comparing hidden states, and featurizers help represent hidden states better. We experiment with several featurizers like crosscoders, Matryoshka SAEs, archetypal SAEs, etc.

Now, it's important to consider the geometry of features. In literature, the Linear Representation Hypothesis (LRH) posits that models represent features as linear directions in activation space. However, recent work has weakened the picture (Engels et al.), as some features live on low-dimensional manifolds, like days of the week arranged on a circle. By design, SAEs take the LRH for granted, which explains why some of their outputs are confusing (besides other failure modes).

More specifically, Bhalla et al. shows that SAEs tile manifolds, spending many directions to cover localities of curved manifolds. We try to recover concept manifolds post-hoc from our trained SAEs by using atom correlations, Ising-style coupling analysis, etc., but without great success.

figure 1

The BSF

Fortunately, the same group recently introduced the block-sparse featurizer (BSF). It's similar to a classic SAE, but its "code" is split into small groups of latents (called blocks) and the sparsity penalty is applied to whole blocks rather than to individual directions. This works if we assume a specific underlying hypothesis behind the data-generating process over activations x∈Rdmodelx \in \mathbb{R}^{d_{model}}: there is a dictionary of GG concepts, where the gg-th occupying a low-dimensional subspace spanned by an orthonormal frame Dg∈Rb×dD_g \in \mathbb{R}^{b \times d} (with b≪db \ll d), and each activation is a sparse sum of contributions from a few of them:

x=∑g∈SzgDg+ε,S⊆{1,…,G},∣S∣≪G,x = \sum_{g \in S} z_g D_g + \varepsilon, \qquad S \subseteq \{1,\dots,G\},\quad |S| \ll G,

where zg∈Rbz_g \in \mathbb{R}^b is the activation's position within concept gg's manifold and SS indexes the concepts present in the input.

We only study the paper's top-kk Grassmannian variant, where the encoder is the decoder's transpose (and they are both orthonormal matrices), and the only learned parameters are the frames D=(D1,…,DG)D = (D_1, \dots, D_G) themselves:

z(x)=Πk(xD⊤),x^=z(x)⋅D,min⁡D∥x−x^∥22,z(x) = \Pi_k\left(x D^\top\right), \qquad \hat{x} = z(x) \cdot D, \qquad \min_D \lVert x - \hat{x} \rVert_2^2,

where Πk\Pi_k zeroes all but the kk blocks of largest norm ∥zg∥2\lVert z_g \rVert_2. This is basically the SAE literature top-kk, but over blocks, not latents. Since Πk\Pi_k sees each block only through its norm, selection is sparse across concepts, but dense within each one, so a whole color circle, for example, can switch on as one atom. This should theoretically reduce manifold tiling.

Contributions

This post does not cover the diffing project, but rather focuses on the BSF and the block paradigm it introduces. More concretely, we:

The Manifold Zoo

Before deploying BSFs on real models, we test them in a setting where we control the data-generating process. The "Manifold Zoo" is a synthetic world of M=128M = 128 features, each with a specific geometry. 6464 of them are straight lines, and the other 6464 are curved manifolds (circles, helices, tori, disks, swiss rolls). Each feature is embedded in a dd-dimensional activation space (d=128d = 128) through a random orthonormal frame. We generate a dataset of 300,000300,000 samples, where every sample is a sparse sum of ∣S∣=4|S| = 4 concepts.

figure 2

In order to test the interpretability of a featurizer like the BSF, we match each feature from the dataset to the one block with the highest firing correlation. We can now score how well that block recovers the feature by measuring the R2R^2 between the block's reconstruction and the feature's true contribution to the sample. Sometimes, we also look at how many blocks are needed for a feature to get R2>0.8R^2 > 0.8.

Unless stated otherwise, the BSF that we train has G=256G = 256 blocks of dimension b=3b = 3 with top-k=4k = 4 selection. Note how the capacity is generous: with 2 blocks per concept, the sparsity is matched to ∣S∣|S|, and dimension bb is matched to the largest dimension a manifold can have. In other words, this should be the perfect setting for the BSF to decode each feature's geometry.

We experiment with two variants of the dataset: one where concepts fire independently (the independent dataset), and one where we introduce correlation between which concepts fire together (the correlated dataset). Feature correlation is a classic pathology for SAE failure, so we consider it necessary to also test this scenario.

BSF vs classic SAE

First, we want to validate that the BSF truly recovers concept geometry better than an SAE. We train an SAE with an equal number of parameters and similar training settings (top-k⋅bk\cdot b, etc.). To make the comparison fair, we also equip the SAE with post-hoc manifold recovery methods, which, after training, group its atoms into candidate manifolds based on their firing statistics (ValuePCA and partial-correlation grouping). As you can see in the table, on the independent dataset the BSF is superior, but post-hoc recovery works quite decently, too. Once we introduce correlation in the dataset, the BSF is untouched, but the SAE loses accuracy. This is expected, as the post-hoc methods rely on correlation statistics to find the manifolds, so concept co-firing may fool concepts into merging.

MethodIndependentCorrelatedSAE (raw atoms)0.360.36SAE + post-hoc (ValuePCA)0.600.41SAE + post-hoc (partial-corr.)0.650.59BSF0.770.77\begin{array}{lcc} \textbf{Method} & \textbf{Independent} & \textbf{Correlated} \\ \hline \text{SAE (raw atoms)} & 0.36 & 0.36 \\ \text{SAE + post-hoc (ValuePCA)} & 0.60 & 0.41 \\ \text{SAE + post-hoc (partial-corr.)} & 0.65 & 0.59 \\ \textbf{BSF} & \textbf{0.77} & \textbf{0.77} \\ \end{array}

BSF consistency across seeds

A big problem with SAEs is that they are not stable, as identical models trained with different seeds can produce significantly different dictionaries (see archetypal SAEs). We test BSFs by training 10 of them on the same data, varying only the seed.

In the table below, we match each feature to the block whose firing correlates most with it, and measure how well that block actually recovers the feature. We observe that segments, spheres, helices, and tori are almost always well recovered by one block, while the other concept manifolds flicker from seed to seed.

figure 3

We now ask how many blocks a feature needs to reach R2>0.8R^2 > 0.8. The answer is either 1, 2, or more than 8 (which we count as not recovered). Interestingly, every concept is captured by exactly one block in at least one of the 10 seeds, suggesting that no feature is intrinsically impossible to recover.

figure 4

Finally, we pick two seeds and compute the correlation between their blocks' firings. The correlation matrix is quite clean, close to an identity. The identity fades for those blocks whose corresponding features are more likely to split across multiple blocks:

figure 5

We also repeat the identical 10-seed experiment on the correlated dataset. The results are basically the same, with similar R2R^2 scores, and similar feature splitting patterns:

figure 6

Testing if hyperparameters cause splitting

So we have shown that BSFs are pretty consistent, but while doing so, we discovered that some features arbitrarily split among multiple blocks, even on the independent dataset, where there is no spurious correlation between features. This is quite reminiscent of how classical SAE directions tile manifolds. In our case, manifolds split among blocks.

Two big factors of tension in any SAE are its capacity (total number of atoms) and its sparsity, and both are known to shape its pathologies (Chanin et al.; Gao et al.). Thus, we wonder if our hyperparameters cause splitting, so we first sweep G∈{128,256,512}G \in \{128, 256, 512\}, but to no avail. Not even G=M=128G = M = 128 works, i.e., having as many blocks as features, where we would have hoped that we could get a feature-block bijection:

figure 7

We also vary sparsity. As our featurizer uses top-kk, we sweep k∈{4,6,8}k \in \{4, 6, 8\}, but that only degrades performance as kk increases, which is expected as ∣S∣=4|S| = 4. For k=8k = 8, not a single feature is captured by one block anymore, as they all split across multiple blocks:

figure 8

So why do splits form?

Let's try and visualize some of the features that split across two blocks. In the figure below, every dot is a sample where the feature is active (we look at one swiss-roll and one helix feature), colored by its position on the manifold. A dot keeps its color in every panel of the row. The two middle columns show what each of the two blocks capture of the feature, and the last column is their sum. We observe that when a manifold splits among two blocks, the blocks tend to be near-orthogonal complements.

figure 9

As previously noted, when a feature splits along multiple blocks, it's almost always exactly 2. For 83 such block pairs, we track the pair through BSF training. The first panel shows the geometric overlap between the two blocks' subspaces, and the second shows each block's firing correlation with the feature (blue for the first block locked onto the feature, and red for the second block):

figure 10

First, see how both blocks start specializing on the same feature around the same time (~epoch 3). For a while, their subspaces grow similar, at ~0.3 overlap (2x of the random baseline, so definitely not copies of one another). But then they pull apart, ending up to be almost orthogonal components of the same feature.

We guess that in the first epochs blocks are learning the feature's strongest axes of variation, and that's why they grow together. After that sharp growth, both blocks have high correlation with the feature, so they are officially "locked" to it, but they are still significantly different. Each block gets gradient from the samples it fires on, so each gets reinforced exactly on the differences it has compared to the other, which drives the blocks to become almost orthogonal. Once they split, each covers variation in samples that the other misses, so gradient descent has no incentive to eliminate one of them.

But we don't fully understand this phenomenon. Given that the pair blocks decompose the manifold orthogonally rather than discovering different regions, they would best reconstruct a sample if they fired together, yet k=∣S∣k = |S| discourages multiple blocks per feature at once, so the pair of blocks tends to not fire together that often (although they do fire together above chance). We also ablate Aux-K (an auxiliary loss standard for top-K SAEs) and k-annealing, and find that none of them significantly contributes to block splitting.

So, just as SAEs tile manifolds, BSFs also split manifolds (although at a smaller rate). We are unsure of the reason, but discover the useful pattern of how block pairs evolve, which we will use in the following section to develop a new BSF architecture.

A solution to feature splitting: Tournament Top-K

A lot of SAE problems like feature composition, hedging, absorption, etc. imply that some SAE latents share directions (i.e., they are not orthogonal). In our case, feature splitting causes the same thing, so it is reasonable to look into known SAE solutions. For example, MP-SAEs decode by greedily subtracting atoms, but this does encourage a new way of feature composition. Ort-SAEs penalize similarities between all pairs of latents, but this is expensive and we also found it to be too discriminatory in the first stages of training. However, we are inspired by Ort-SAEs to propose Tournament Top-K:

For each sample, go through the blocks in decreasing order of magnitude ∣zg∣|z_g|, keeping a running set BB of accepted blocks. For the next block gg, if some h∈Bh \in B overlaps with its geometry sufficiently (we say Ωgh>0.1\Omega_{gh} > 0.1), then we skip gg and count a win for hh. Otherwise, add gg to BB. We stop at ∣B∣=k|B| = k (the BSF's sparsity).

Note that blocks gg and hh could meet again in a future duel. However, if at some point Ωgh>0.5\Omega_{gh} > 0.5, the block with more wins will permanently win the feature, which means that in any other future duel we will keep the winning block in BB, and exclude the other.

To understand why we are using this 2-threshold system, consider this graph: (note that Ω\Omega is a different metric for geometric similarity than the one used above, comparing only directions blocks actually use, thus we have a different scale):

figure 11

In the first epochs, there are many pairs of blocks that overlap, yet they naturally set on different features later (orange line). However, they are initially indistinguishable from blocks that actually split features (blue line). Thus, it's hard to set just one threshold. Our initial 0.10.1 threshold catches true splits about just a third of the time. That's why we make it reversible (a winner is not decided yet). Also for this reason, we only begin dueling after a warmup of ~16 epochs, allowing blocks to specialize a bit.

Note that the 0.50.5 threshold is naturally uncrossable by pairs of blocks. However, our 0.10.1 duels push blocks up there: if two blocks really share a feature, they get chosen in an alternate manner by the 0.10.1 duel (depending on the sample), which pushes them both to reconstruct the whole feature alone. The overlap thus climbs past 0.50.5; if two blocks don't actually share a feature, the 0.10.1 duel exclusivity doesn't actually do anything, so each block just develops on its own, and they never meet again in a duel.

We test on the correlated dataset, which is our hardest setting. Tournament Top-K raises mean recovery R2R^2 significantly, and also lifts the fraction of well-recovered manifolds from 0.770.77 to 0.900.90, significantly cutting the splitting phenomenon:

figure 12

Varied Dim Blocks

We also note that feature manifolds have varied dimensionalities, yet all blocks are 3D. This leads to feature merging: for example, we observed cases of up to three linear features being explained by one block. We thus give our BSF a mix of block sizes (128 x 1D + 38 x 2D + 90 x 3D blocks), and test whether training self-sorts. We try the following regimes (and fail in each of them):

We measured that a BSF with a mix of block sizes would reach a smaller loss than a regular BSF if it learned how to map features to blocks optimally, yet it seems training can't reach this. Once a feature sticks to a block in the first epochs, moving it is impossible for gradient descent, because a block move raises the loss for a significant time. We think interventions in the spirit of Tournament Top-K might be needed, but we leave that as a homework for the reader.

Applying BSFs on real models

We have also trained BSFs on real models (DINO and CLIP). We sweeped through many configurations to conclude our best setting: 16,38416,384-dimensional dictionaries, with top-K=64K = 64, and blocks of dimension 33. We also train Matryoshka SAEs at the same scale. Some of the experiments we run with the BSFs:

Note that this section is very brief. The point of this post is to study the BSF itself, while the aforementioned experiments diverge from that. Many of these will be further detailed in the actual model diffing paper.

Block Crosscoders

Crosscoders are a really important mechanistic interpretability tool for model diffing, and we wonder if the block paradigm extends to them. In this experiment, we maintain the top-K method. Consider a paired sample (xA,xB)(x^A, x^B) of hidden states from some layers in models AA and BB. We get the code by doing:

z=Πk(xAEA+xBEB),z = \Pi_k\left(x^A E^A + x^B E^B\right),

where Πk\Pi_k keeps the kk blocks of largest norm ∣zg∣2|z_g|_2, exactly as before. The code then decodes through two dictionaries,

x^m=∑g∈ΠkzgDgm,m∈A,B,min⁡D∑m∣xm−x^m∣22,\hat{x}^m = \sum_{g \in \Pi_k} z_g D_g^m, \quad m \in {A, B}, \qquad \min_{D} \sum_m |x^m - \hat{x}^m|_2^2,

where Πk\Pi_k selects the kk blocks of largest joint norm.

Merging Problem

We do notice a structural problem with this simple extension of the crosscoder though. Testing it on a new Manifold Zoo, with 64 shared features, 32 exclusive to AA, and 32 exclusive to BB, we find that the crosscoder often assigns an AA-only feature to block gg, while also assigning a BB-only feature to the same block gg.

Note how this problem is specific to the block-crosscoder. In a normal crosscoder, a latent (a dimension of zz) is a single scalar, so both decoders multiply by the same scalar. Thus, the two concepts would have to fire together with same intensity at all time. In a block however, the corresponding zgz_g has bb dimensions, so one decoder can use some of them, while the other decoder uses the rest. This is especially lucrative when the sum of the dimensionalities of the A-only and B-only features is ≤b\le b. That way, the two concepts can stay fully independent within the same block.

Dedicated Feature Columns

To fix this, we use dedicated feature columns (DFCs, by Jiralerspong et al.) from the crosscoder literature. This method pre-partitions blocks into 3 pools: shared, A-only, and B-only. So basically DAD^A is zeroed-out on the B-only blocks, and vice versa. On our toy dataset, feature collisions become much rarer with block-DFC crosscoders.

Results

We train a block-DFC crosscoder on paired DINO/CLIP activations at layer 10, with G=8192G = 8192 blocks, and a pool distribution of 4096 shared blocks, 2048 DINO-only, and 2048 CLIP-only. EV's are healthy at 0.71 for DINO and 0.66 for CLIP, and we only have 1 dead block.

We now match each single-model BSF's blocks with the crosscoder's blocks, by firing rate correlation:

DINO BSFCLIP BSFblocks recovered in the crosscoder82%53%pool of recovered blocks (shared / DINO / CLIP)26% / 73% / 1%62% / 5% / 33%\begin{array}{lcc} \hline & \textbf{DINO BSF} & \textbf{CLIP BSF} \\ \hline \text{blocks recovered in the crosscoder} & \mathbf{82\%} & 53\% \\ \text{pool of recovered blocks (shared / DINO / CLIP)} & 26\% \,/\, 73\% \,/\, 1\% & 62\% \,/\, 5\% \,/\, 33\% \\ \hline \end{array}

It looks like the crosscoder really learns DINO's dictionary and leverages its exclusive pool, but treats CLIP as mostly shared and does not even cover it well. It's an open question whether CLIP's visual features are a subset of DINO, or CLIP's code is just harder to recover.

Conclusion

The BSF delivers on its core promise: when features live on manifolds, it recovers them as single units, while regular SAEs tile them into fragments. Still, it does not escape the classic SAE pathologies, as features keep splitting across blocks even with no correlations in the data. The splitting has structure though (split blocks end up near-orthogonal complements of each other), and our Tournament Top-K exploits this geometry to recover much of what vanilla top-k loses. Mixing block dimensionalities fails, with features locking into wrong-sized blocks early in training. Extending blocks to crosscoders brings a merging problem, which dedicated feature columns partly fix. Overall, we think the BSF is a step towards the right direction, i.e., developing featurizers suited to the data generating process. We are going to use the BSF intensely in our follow-up model diffing paper.

The code for this project is available at github.com/lolismek/bsf-analysis. If you use or build on this work, please cite it (an arXiv version is coming soon).