Alex's mech-interp place

by Alex Jerpelea

These are my mechanistic interpretability projects from the summer of 2026, done either in labs at Columbia University or on my own. Each box opens the write-up of one project.

Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry

August 2026

We plant a controllable latent variable inside natural-looking text via subliminal steering. A small transformer trained on the corpus tracks the Bayesian posterior over the variable, and arranges its 8 states on a ring, in the exact order of the Markov chain. This gives stronger evidence that belief states form beyond toy synthetic data, and empirically ties them to the geometry of concepts.

A deeper analysis of Goodfire's Block Sparse Featurizers

August 2026

We stress-test the block-sparse featurizer, an SAE whose atomic unit is a small subspace rather than a single direction, on a synthetic manifold zoo and on real vision models. BSFs recover manifold features far better than classic SAEs but still split them across blocks; we characterize the split geometry, propose Tournament Top-K as a fix, and extend the block paradigm to crosscoders.

The Emotion Circumplex in a Multilingual LLM's Activation Space

August 2026 · from the Mosaic-Emo paper, under review

For ~2,300 self-reported affective states across eight languages, we extract activation centroids from Qwen3-8B and ask whether Russell's valence–arousal circumplex organizes them. The circumplex turns out to be the leading geometry even of the broader, open-vocabulary states, stays present but entangled in the weaker languages, and transfers across languages: one language's plane recovers every other language's circle, with Mandarin the best source for almost every target.

Multi-agent systems

not really mech-interp, but studying multi-agent systems and their failure modes is essential for safe agentic AI

Shared Epistemic State in Multi-Agent Systems

August 2026 · work in progress

Messages between agents are lossy compressions of their epistemic state: beliefs, assumptions, and confidence that never fully reach the other agents. Case studies of MAS failures on GAIA, one controlled experiment for each type of information asymmetry, two belief-sharing mechanisms evaluated across relay and hub topologies (a null result so far), and ongoing work on reading the epistemic state directly from activations.

Other

small experiments and other work

Reading list

129 papers · updated August 2026

Everything I consumed to learn mech interp (the ARENA course, learnmechinterp.com, and all the interp papers from my research notes), plus the multi-agent systems literature behind the Shared Epistemic State project.

J-Think: soft thinking only in the workspace

July 2026

Latent chain-of-thought that never leaves the middle of the model: the thinking loop runs only through Qwen's workspace layers, with Jacobian-lens linear bridges skipping the sensory and motor layers entirely. The mechanics work and give a small lift over direct answering, but the lift doesn't grow with more steps.

2D transformers: reading the depth axis

June 2026

A second small transformer is trained jointly with the backbone to read the ladder of hidden states (the depth axis) as a sequence and produce the readout, hoping the backbone bends its space to be read. Six experiments in, the depth-reader ties but never beats the top-layer readout: the residual stream already telescopes the whole ladder into the top state.

Steering without the layer sweep

June 2026

Early-stage: activation steering where an attention mechanism decides how the steering vector is distributed across layers, so you never grid-sweep for the one right injection point.