← Alex's mech-interp place

Shared Epistemic State in Multi-Agent Systems

by Alex Jerpelea · August 27, 2026 · joint work with Yusen Zhang (Columbia DAP Lab) · code

Abstract

Multi-agent systems (MAS) with complex communication graphs are notoriously hard to scale: adding more agents does not monotonically improve performance. We identify and characterize this failure mode, which we term lost-in-propagation: as the number of agents grows, task performance first rises and then declines. A natural hypothesis is that agents simply lack sufficient shared context. We argue this is not the root cause: naively broadening context sharing does not resolve the degradation, because raw context is unbounded and increasingly dilutes the signal each agent must act on. Instead, we believe the missing ingredient is epistemic status: each agent's explicit account of what it concluded, how confident it is, and on what basis. We build mechanisms for agents to produce and exchange this epistemic status rather than raw content, giving each agent a compact, structured view of its collaborators' beliefs and uncertainty.

This is a write-up of work in progress. The diagnosis is well supported by case studies and controlled experiments, but our first lightweight belief-sharing mechanisms do not yet significantly beat vanilla baselines, which is pushing us toward reading the epistemic state directly out of the model instead of asking for it in text.

Case studies: why multi-agent systems fail

Multi-agent collaboration has become a core method for extending the performance of agents in complex reasoning, such as coding, math, and agentic tasks. However, multi-agent communication is reported to have diverse problems, such as miscommunication and misalignment. The MAST taxonomy (Cemri et al.) sorts these failures into system design issues, inter-agent misalignment, and task verification failures.

As a first step, we ran two canonical MAS, ChatDev and Magentic-One (a hub-and-spoke system for web research, where an Orchestrator holds a "ledger" of collected facts and current plan, and delegates one step at a time to four specialists), on 15 GAIA tasks each, and hand-coded every trace with the MAST failure modes. Inter-agent misalignment failures were by far the most common:

Table of MAST inter-agent misalignment failure counts for ChatDev and Magentic-One on 15 GAIA tasks each
MAST inter-agent misalignment failure modes counted over 15 GAIA runs of ChatDev and 15 of Magentic-One. "Ignored other agent's input" and "information withholding" are near-ubiquitous.

Three Magentic-One case studies stuck with us:

Does not know its collaborator's status. The WebSurfer can't open YouTube transcripts, but the Orchestrator's plan depends on extracting the video transcript. Although it is obvious that the WebSurfer can't do that action, the Orchestrator repeats the same transcript-extraction instruction ~8 times.

Does not know the answer is ready. The WebSurfer finds the desired blog post and shows it to the Orchestrator, but the Orchestrator ignores it and does not click it, never acknowledging that it received exactly what it wanted from another agent. It proceeds to loop on re-making the plan.

Does not know it is wrong. The WebSurfer brings forward a fake answer, and the Orchestrator knows it. But running out of turns, the Orchestrator "panics" and decides to use the fake answer anyway.

We saw similar problems in the decentralized setting, running AutoGen group chat on GAIA:

In each case, some agent held the crucial piece of information (I can't do this, we're done, this is fake), and it never made it across to the other agents. These are failures of first-order theory of mind between agents.

Framing the problem

When agent A communicates with agent B, its message is generated from an underlying epistemic state (beliefs, assumptions, and confidence that A holds but does not necessarily write down). We model this state as a latent variable Z, with the transmitted message M = f(Z). The verbalization function f is lossy: M under-determines Z, so parts of A's state never reach B.

Diagram: agent A's epistemic state Z is verbalized through a lossy function f into the message M sent to agent B
The message M is a lossy verbalization of A's epistemic state Z.

Each multi-agent graph can be abstracted to a hub or a relay, and in both cases agents communicate with the output and leave the working trace inside the model:

Diagram: multi-agent graphs abstracted into hub and relay topologies, with working traces hidden inside each agent
Multi-agent graphs abstract to hub or relay topologies. Agents exchange outputs; the working trace stays inside each agent.

So the agents' contexts are not fully shared, leading to different interpretations:

Existing work mostly attacks the explicit side by sharing more raw context, but raw context is unbounded and increasingly dilutes the signal each agent must act on. Before building anything, we designed one controlled experiment for each type of asymmetry.

Two controlled experiments

The explicit version: lost-in-propagation

Laban et al. show that LLMs get lost in multi-turn conversation; lost-in-propagation is the multi-agent version of it. The question, for the explicit asymmetry, is whether splitting the work across more agents helps or hurts, given the same total budget of tokens.

We test this with a relay chain of agents, with the same total token budget at every chain length N, on FanOutQA, a benchmark that requires discovering multiple facts. "Fact recall" measures what percentage of the facts was recovered, while exact match is more strict on the final answer.

Plot: exact match, fact recall, and completion tokens versus relay chain length on FanOutQA at a fixed 32k token budget
Accuracy and token spend vs chain length (32k total budget, 40 FanOutQA tasks). Exact match peaks at N = 2 and then declines; fact recall collapses by N = 8 while completion tokens climb toward the budget cap.

Performance first rises (two agents beat one) and then declines, at constant total context. Each hand-off passes through the lossy f, and the losses compound.

The implicit version: belief injection

To isolate the implicit asymmetry, we inject beliefs into the system prompts of a relay of 3 agents, hidden from the other agents, in three variants: probe (each agent gets 3 different beliefs), homo (each agent gets the same 3 beliefs), and none (no belief injection). We evaluate on both FEVER and MATH level-5. Belief examples:

"I find substituting a fresh variable for a messy radical expression far more elegant than blind algebraic grinding." (MATH)

"Most striking claims that circulate in claim-verification sets are distortions; your default posture is disbelief unless the fact is solidly familiar to you." (FEVER)

Bar chart: accuracy of none, homo, and probe belief-injection variants on MATH level-5 and FEVER
Belief injection in a 3-agent relay, on MATH level-5 and FEVER. On FEVER, shared beliefs (homo) beat no beliefs, and divergent hidden beliefs (probe) fall below both.

FEVER is the headline, with homo > none > probe. Agents holding the same hidden beliefs outperform belief-free agents, and agents holding different hidden beliefs underperform them. On MATH-L5 the differences are not significant, which could be attributed to the fact that MATH-L5 tasks are much more objective, so beliefs have less influence.[2]2.For the FEVER run we also made the injected beliefs stronger. So hidden belief divergence alone can move task accuracy.

Proposed solutions: a belief table and a belief board

The case studies point at weak inter-agent theory of mind, so the belief system is about making first-order ToM easier: instead of B having to infer what A thinks, A just tells B its opinions directly. We built two MAS-agnostic v0 prototypes.

The belief table reconstructs, from A's trace, a table of A's beliefs about a set of salient objects. Some of those objects are objective (a Wikipedia fact), where A's "belief" is really just a found fact, acting like shared memory. Others are subjective ("is this task even do-able?"), where A's belief is a genuine stance: if A loses hope near the end, that warns B its answer is probably a hallucination. An auxiliary observer LLM reads A's full trace at each hand-off and writes a short, revisable ledger of typed entries (observation vs. belief), each with an object, a claim, a confidence, and an author.

Diagram of the belief table: an observer LLM maps agent A's trace to typed entries with object, claim, confidence, author
The belief table: an observer LLM reads A's full trace (including hidden reasoning) and maps it to typed entries. Objective objects act like shared memory; subjective objects carry stances like hallucination warnings.

The belief board is the simpler, more direct version: instead of a second model reconstructing A's beliefs after the fact, we hand A the tools to jot down its beliefs itself, live, as it works. Whenever A establishes a fact, computes a value, or rules out a hypothesis, it calls add_belief or revise_belief and writes [object, claim, confidence] to a shared board, attributed to A. Nothing is ever deleted: a revision appends a corrected entry pointing back at the old one, so B sees not only what A currently believes, but that it changed its mind. The board rides alongside the normal hand-off as a lateral channel, and every agent sees it rendered at the start of its turn. It's free-er and fuzzier than the observer table: A decides on its own what's worth noting, which is lighter-weight (no extra model call) but leaves the object set unstructured.

Diagram of the belief board: agent A appends and revises belief entries on a shared board via tools
The belief board: A writes beliefs live with its own tools; revisions append and point back, so B sees the change of mind.

Evaluation

Designing a fair evaluation setup took a few iterations.[3]3.We first tried recreating the working-memory baselines of the G-Memory and LatentMem papers (MemoryBank, Generative, Voyager-mem, ChatDev-mem, MetaGPT-mem), and found the implementations are not faithful to the actual papers; it is genuinely hard to make a MAS-agnostic working memory, which is probably why most papers focus on external memory. Also, most canonical benchmarks are solvable by single agents, so multi-agent becomes awkward. First, the MAS should be such that each agent has an internal ReAct loop before communicating with other agents; this seems trivial, but most MAS papers actually don't respect it (in MacNet or DyLAN, all nodes discuss one agentic turn, so the nodes are not really agents, and beliefs become redundant). Second, our system relies on information asymmetry, so we develop two simplified MAS that reflect the two types of it: temporal (Relay) and spatial (Hub).

Relay is a chain of agents passing notes to one another. Only a free-text hand-off note crosses each edge, never the predecessor's transcript; each shift re-reads the task plus that note. The objective is to catch temporal asymmetry: errors arising when an agent passes task status mid-execution to another agent.

Diagram of the Relay topology: three shifts work one task in sequence, passing only a hand-off note
Relay: temporal asymmetry. K = 3 shifts work one task in sequence, each with fresh context, a budget, tools, and an internal ReAct loop.

Hub is a star topology similar to AutoGen group chat, simplified so the orchestrator decomposes the task into subtasks once and then aggregates all worker reports into one answer. Each worker sees the full task but none of the others; only the orchestrator integrates. The point is to catch topology-induced asymmetries more cleanly.

Diagram of the Hub topology: an orchestrator decomposes the task for blind workers and merges their reports
Hub: spatial asymmetry. Blind workers report back to a tool-less orchestrator that decomposes, merges, and commits.

Because our belief board targets ToM issues rather than factual recall, the appropriate baselines are also ToM-flavored rather than full factual memory: full (just keep all logs, for full transparency), sop (structured verdict / evidence / next-steps hand-offs), down (debate on demand), extract (an observer-style extractor), and board_inert (a tool-use control for the board). Each arm runs on both topologies, over GAIA, PDDL, FEVER-compound, FanOutQA, and GPQA-Diamond as a closed-book null anchor with no information asymmetry.

Table 1: task outcome accuracy across all arms and benchmarks on Relay and Hub, Qwen3.6-35B-A3B; no arm-vanilla difference is significant

Across the full grid, no arm beats vanilla significantly, our board included. There was a lot of debugging and tuning to do, such that the MAS failures are genuine and not just artifacts of bad implementation; the setup now produces actual information asymmetry, uniformly comparable across baselines, and the belief systems mechanically work. But at this scale, lightweight text-level belief sharing does not yet move the needle.

Ongoing work: reading Z directly

A possible explanation for the null result is that the treatment is applied at the wrong level: if f is lossy, any mechanism that asks A to verbalize its state is itself downstream of f. So we are now searching for methods that expose more of Z without going through the message channel.

Probing with random questions. The message M captures only a part of A's state (the region f−1(M)). To reach the rest, we make B ask A a couple of questions q1…qn that are deliberately unrelated to the task. Because A answers from inside its working context, every answer is conditioned on Z, so off-topic answers can leak information from the part of Z that M left unexposed, and B can do Bayesian inference on them. The questions must be unrelated: related questions would be, in a way, conditioned on M (that's the whole context B has), and we want to discover information outside of M.

Diagram: random task-unrelated questions from B probe parts of A's state Z that the message M left unexposed
Random, task-unrelated questions can land anywhere in Z, including the part left unexposed by M.

Reading Z with the J-lens. The J-lens (Anthropic, 2026) reads A's global workspace: the set of concepts A is silently entertaining at a given moment (things A could answer questions about if we paused it and asked, but which it does not necessarily write down). This is exactly the kind of content we mean by Z: beliefs and hypotheses that are active and reportable, yet absent from M. The lens yields a readout g(Z), the top-k workspace tokens at a middle layer, including never-verbalized ones, which we compress into a short structured blurb and pass to B alongside M, bypassing the lossy f entirely.

Diagram: the J-lens reads a workspace readout g(Z) directly from A's hidden states, bypassing the lossy verbalization f
The J-lens readout g(Z) is taken from A's hidden states and does not pass through the lossy f.

Conclusion

The case studies, the lost-in-propagation curve, and the belief-injection experiment all point at the same bottleneck: messages are lossy compressions of epistemic state, so parts of what an agent knows never reach its collaborators. Our first treatments, the text-level belief table and the belief board, do not yet significantly beat vanilla baselines. The open question, and where the project is headed, is at what level epistemic-state sharing has to happen: text, tools, or activations.

  1. In one run, an agent thinks that basil is not a vegetable, and it convinces the whole MAS of this fact.↩
  2. For the FEVER run we also made the injected beliefs stronger.↩
  3. We first tried recreating the working-memory baselines of the G-Memory and LatentMem papers (MemoryBank, Generative, Voyager-mem, ChatDev-mem, MetaGPT-mem), and found the implementations are not faithful to the actual papers; it is genuinely hard to make a MAS-agnostic working memory, which is probably why most papers focus on external memory. Also, most canonical benchmarks are solvable by single agents, so multi-agent becomes awkward.↩

The code for this project is available at github.com/lolismek/mas-research. This is ongoing joint work within Columbia's DAP Lab with Yusen Zhang; the figures come from our internal progress slides. If you use or build on this work, please cite it.