← back to the mech-interp place
Reading list
by Alex Jerpelea
Everything I consumed to learn mechanistic interpretability, plus the multi-agent systems literature behind the Shared Epistemic State project, extracted from my research notes, roughly in chronological order.
Start here
Mechanistic interpretability
- Toy Models of Superposition — Nelson Elhage et al., 2022
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small — Kevin Wang et al., 2022
- Discovering Latent Knowledge in Language Models Without Supervision — Collin Burns et al., 2022
- Steering GPT-2-XL by Adding an Activation Vector — Alexander Turner et al., 2023
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets — Samuel Marks et al., 2023
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning — Trenton Bricken et al., 2023
- The Linear Representation Hypothesis and the Geometry of Large Language Models — Kiho Park et al., 2023
- Function Vectors in Large Language Models — Eric Todd et al., 2024
- Improving Dictionary Learning with Gated Sparse Autoencoders — Senthooran Rajamanoharan et al., 2024
- Not All Language Model Features Are One-Dimensionally Linear — Joshua Engels et al., 2024
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet — Adly Templeton et al., 2024
- Transformers represent belief state geometry in their residual stream — Adam Shai et al., 2024
- The Geometry of Categorical and Hierarchical Concepts in Large Language Models — Kiho Park et al., 2024
- Transcoders Find Interpretable LLM Feature Circuits — Jacob Dunefsky et al., 2024
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 — Tom Lieberum et al., 2024
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders — David Chanin et al., 2024
- Sparse Crosscoders for Cross-Layer Features and Model Diffing — Jack Lindsey et al., 2024
- LatentQA: Teaching LLMs to Decode Activations Into Natural Language — Alexander Pan et al., 2024
- ICLR: In-Context Learning of Representations — Core Francisco Park et al., 2025
- Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models — Thomas Fel et al., 2025
- Constrained belief updates explain geometric structures in transformer representations — Mateusz Piotrowski et al., 2025
- Detecting Strategic Deception Using Linear Probes — Nicholas Goldowsky-Dill et al., 2025
- Insights on Crosscoder Model Diffing — Siddharth Mishra-Sharma et al., 2025
- Sparse Autoencoders Do Not Find Canonical Units of Analysis — Patrick Leask et al., 2025
- Circuit Tracing: Revealing Computational Graphs in Language Models — Emmanuel Ameisen et al., 2025
- On the Biology of a Large Language Model — Jack Lindsey et al., 2025
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability — Adam Karvonen et al., 2025
- Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders — David Chanin et al., 2025
- The Origins of Representation Manifolds in Large Language Models — Alexander Modell et al., 2025
- Detecting High-Stakes Interactions with Activation Probes — Alex McKenzie et al., 2025
- From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit — Valérie Costa et al., 2025
- Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures — Mark Muchane et al., 2025
- SPARC: Concept-Aligned Sparse Autoencoders for Cross-Model and Cross-Modal Interpretability — Ali Nasiri-Sarvi et al., 2025
- OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features — Anton Korznikov et al., 2025
- Eliciting Secret Knowledge from Language Models — Bartosz Cywiński et al., 2025
- Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry — Thomas Fel et al., 2025
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — Adam Karvonen et al., 2025
- The Bayesian Geometry of Transformer Attention — Naman Aggarwal et al., 2025
- Verbalizable Representations Form a Global Workspace in Language Models — Wes Gurnee et al., 2026
- Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs — Thomas Jiralerspong et al., 2026
- From Atoms to Trees: Building a Structured Feature Forest with Hierarchical Sparse Autoencoders — Yifan Luo et al., 2026
- Symmetry in language statistics shapes the geometry of model representations — Dhruva Karkada et al., 2026
- SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data — David Chanin et al., 2026
- The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors — Raphael Sarfati et al., 2026
- Transformers learn factored representations — Adam Shai et al., 2026
- Do Sparse Autoencoders Capture Concept Manifolds? — Usha Bhalla et al., 2026
- Finding Belief Geometries with Sparse Autoencoders — Matthew Levinson, 2026
- MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents — Matthew Levinson, 2026
- Subliminal Steering: Stronger Encoding of Hidden Signals — George Morgulis et al., 2026
- Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence — Andres Nava et al., 2026
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior — Daniel Wurgaft et al., 2026
- SMIXAE: Towards Unsupervised Manifold Discovery in Language Models — Collin Francel, 2026
- Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space — Eric J. Bigelow et al., 2026
- Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders — Tue Cao et al., 2026
- Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds — Thomas Fel et al., 2026
- Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability — Seyed Arshan Dalili et al., 2026
Multi-agent systems
- Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate — Kai Xiong et al., 2023
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate — Tian Liang et al., 2023
- A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration — Zijun Liu et al., 2023
- Let Models Speak Ciphers: Multiagent Debate through Embeddings — Chau Pham et al., 2023
- Theory of Mind for Multi-Agent Collaboration via Large Language Models — Huao Li et al., 2023
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions — Hyunwoo Kim et al., 2023
- Multi-agent Reinforcement Learning: A Comprehensive Survey — Dom Huh et al., 2023
- Language Agents as Optimizable Graphs — Mingchen Zhuge et al., 2024
- Belief sharing: a blessing or a curse — Ozan Catal et al., 2024
- When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn Settings — Jérémy Perez et al., 2024
- Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models — Logan Cross et al., 2024
- Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems — Guibin Zhang et al., 2024
- G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks — Guibin Zhang et al., 2024
- A Scalable Communication Protocol for Networks of Large Language Models — Samuele Marro et al., 2024
- DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving — Yuhan Liu et al., 2024
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks — Adam Fourney et al., 2024
- SRMT: Shared Memory for Multi-agent Lifelong Pathfinding — Alsu Sagirova et al., 2025
- Communicating Activations Between Language Model Agents — Vignav Ramesh et al., 2025
- Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection — Bo Yang et al., 2025
- Multi-agent Architecture Search via Agentic Supernet — Guibin Zhang et al., 2025
- SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning — Wanjia Zhao et al., 2025
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025
- LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative Planning — Hanqing Yang et al., 2025
- AutoToM: Scaling Model-based Mental Inference via Automated Agent Modeling — Zhining Zhang et al., 2025
- MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning — Chanwoo Park et al., 2025
- Why Do Multi-Agent LLM Systems Fail? — Mert Cemri et al., 2025
- Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning — Sugyeong Eo et al., 2025
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara et al., 2025
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems — Shaokun Zhang et al., 2025
- Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space — Zhen Zhang et al., 2025
- Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs — Hexiang Tan et al., 2025
- Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control — Alireza Rezazadeh et al., 2025
- G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems — Guibin Zhang et al., 2025
- From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium — Xie Yi et al., 2025
- MIRIX: Multi-Agent Memory System for LLM-Based Agents — Yu Wang et al., 2025
- RCR-Router: Efficient Role-Aware Context Routing for Multi-Agent LLM Systems with Structured Memory — Jun Liu et al., 2025
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning — Sikuan Yan et al., 2025
- Collaborative Belief Reasoning with LLMs for Efficient Multi-Agent Collaboration — Zhimin Wang et al., 2025
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents — Guibin Zhang et al., 2025
- Mem-α: Learning Memory Construction via Reinforcement Learning — Yu Wang et al., 2025
- LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science — Alireza Salemi et al., 2025
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models — Tianyu Fu et al., 2025
- KVComm: Enabling Efficient LLM Communication through Selective KV Sharing — Xiangyu Shi et al., 2025
- LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation — Dongge Han et al., 2025
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts — Yeskendir Koishekenov et al., 2025
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems — Hancheng Ye et al., 2025
- Enabling Agents to Communicate Entirely in Latent Space — Zhuoyun Du et al., 2025
- Ask WhAI: Probing Belief Formation in Role-Primed LLM Agents — Keith Moore et al., 2025
- Latent Collaboration in Multi-Agent Systems — Jiaru Zou et al., 2025
- MemEvolve: Meta-Evolution of Agent Memory Systems — Guibin Zhang et al., 2025
- Recursive Language Models — Alex L. Zhang et al., 2025
- FlashMem: Distilling Intrinsic Latent Memory via Computation Reuse — Yubo Hou et al., 2026
- Toward Efficient Agents: Memory, Tool learning, and Planning — Xiaofang Yang et al., 2026
- Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems — Ruiwen Zhou et al., 2026
- LatentMem: Customizing Latent Memory for Multi-Agent Systems — Muxin Fu et al., 2026
- When Does Multi-Agent Collaboration Help? An Entropy Perspective — Yuxuan Zhao et al., 2026
- Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems — Adam Kostka et al., 2026
- Adaptive Theory of Mind for LLM-based Multi-Agent Coordination — Chunjiang Mu et al., 2026
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents — Ling Yue et al., 2026
- Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents — Nicholas Edwards et al., 2026
- Detecting Multi-Agent Collusion Through Multi-Agent Interpretability — Aaron Rose et al., 2026
- CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery — Ao Qu et al., 2026
- Learning to Interrupt in Language-based Multi-agent Communication — Danqing Wang et al., 2026
- Prism: An Evolutionary Memory Substrate for Multi-Agent Open-Ended Discovery — Suyash Mishra, 2026
- Recursive Multi-Agent Systems — Jiaru Zou et al., 2026
- TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems — Chen Xu et al., 2026
- What Do Agents Communicate? Characterizing Information Exchange in Multi-Agent Systems — Yong Jin Chun et al., 2026
- Multi-agent Collaboration with State Management — Mengyang Liu et al., 2026
- Dynamic Mixture of Latent Memories for Self-Evolving Agents — Dianzhi Yu et al., 2026
- Scaling Behavior of Single LLM-Driven Multi-Agent Systems — Jialing Li et al., 2026
- What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems — Chen Huang et al., 2026
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents — Aofan Yu et al., 2026
- Decentralized Multi-Agent Systems with Shared Context — Yuzhen Mao et al., 2026