These are my mechanistic interpretability projects from the summer of 2026, done either in labs at Columbia University or on my own. Each box opens the write-up of one project.
We plant a controllable latent variable inside natural-looking text via subliminal steering. A small transformer trained on the corpus tracks the Bayesian posterior over the variable, and arranges its 8 states on a ring, in the exact order of the Markov chain. This gives stronger evidence that belief states form beyond toy synthetic data, and empirically ties them to the geometry of concepts.
We stress-test the block-sparse featurizer, an SAE whose atomic unit is a small subspace rather than a single direction, on a synthetic manifold zoo and on real vision models. BSFs recover manifold features far better than classic SAEs but still split them across blocks; we characterize the split geometry, propose Tournament Top-K as a fix, and extend the block paradigm to crosscoders.
For ~2,300 self-reported affective states across eight languages, we extract activation centroids from Qwen3-8B and ask whether Russell's valence–arousal circumplex organizes them. The circumplex turns out to be the leading geometry even of the broader, open-vocabulary states, stays present but entangled in the weaker languages, and transfers across languages: one language's plane recovers every other language's circle, with Mandarin the best source for almost every target.
not really mech-interp, but studying multi-agent systems and their failure modes is essential for safe agentic AI
Messages between agents are lossy compressions of their epistemic state: beliefs, assumptions, and confidence that never fully reach the other agents. Case studies of MAS failures on GAIA, one controlled experiment for each type of information asymmetry, two belief-sharing mechanisms evaluated across relay and hub topologies (a null result so far), and ongoing work on reading the epistemic state directly from activations.
small experiments and other work
Everything I consumed to learn mech interp (the ARENA course, learnmechinterp.com, and all the interp papers from my research notes), plus the multi-agent systems literature behind the Shared Epistemic State project.
Latent chain-of-thought that never leaves the middle of the model: the thinking loop runs only through Qwen's workspace layers, with Jacobian-lens linear bridges skipping the sensory and motor layers entirely. The mechanics work and give a small lift over direct answering, but the lift doesn't grow with more steps.
A second small transformer is trained jointly with the backbone to read the ladder of hidden states (the depth axis) as a sequence and produce the readout, hoping the backbone bends its space to be read. Six experiments in, the depth-reader ties but never beats the top-layer readout: the residual stream already telescopes the whole ladder into the top state.
Early-stage: activation steering where an attention mechanism decides how the steering vector is distributed across layers, so you never grid-sweep for the one right injection point.