← back to the mech-interp place
A small experiment: latent chain-of-thought that never leaves the middle of the model. The recurrent loop runs only through Qwen's workspace layers, and the Jacobian lens lets us skip the sensory and motor layers entirely with three matrix multiplies.
"Verbalizable Representations Form a Global Workspace in Language Models" argues that a transformer's layers split into three regions: early sensory layers that parse tokens into representations, a long middle workspace where those representations live in one shared, verbalizable geometry, and late motor layers that convert the state back into a next-token distribution.
Now look at what thinking costs. In chain of thought, every reasoning step traverses all layers and then collapses into a single discrete token, which re-enters at the bottom. Soft-thinking methods (Coconut and friends) drop the token (the final hidden state is fed back in as the next input embedding), but each thought still makes the full round trip through the motor and sensory layers on every loop. If thinking happens in the workspace, both trips are overhead.
J-Think keeps the loop inside the workspace. On Qwen3-4B (36 layers; workspace ≈ layers 6–28) the recurrent state exits at the top of the workspace and re-enters at its bottom, with everything outside replaced by linear maps: the Jacobian lens already gives a linear approximation of all the layers above the exit, and a small ridge-fit map stands in for the sensory layers below the entry. One latent thought = one pass through the 23 workspace layers plus three matrix multiplies: no token, no motor layers, no sensory layers.
Before any of that, the workspace has to exist in Qwen. Following the paper, I computed the linear CKA between the J-lens vector geometries of every pair of layers. The block structure replicates cleanly on Qwen3.5-4B (below) and Qwen3.6-27B: a small sensory block (layers 0–2), one long high-CKA workspace block spanning nearly all of the depth, and the motor stage at the very top. One difference from the paper's result on Claude Sonnet 4.5: Qwen's sensory block is proportionally much smaller (roughly the first 9% of depth vs. the first third).
Write \(h_{28}\) for the last-position hidden state at the workspace exit (layer 28). A vanilla soft-thinking step would run layers 29–35, unembed, re-embed, and run layers 0–5 again before the next thought. J-Think replaces that whole detour with linear bridges. First, the Jacobian-lens matrix at layer 28 stands in for the motor layers, predicting the final-layer state directly:
$$\hat h_{\text{fin}} \;=\; \mu_{\text{fin}} \;+\; J^{\text{mot}}\,\big(h_{28} - \mu_{\text{mot}}\big)$$Second, the predicted final state is mapped back to embedding space. \(W_a\) is the least-squares map from unembedding to embedding rows, and because Qwen3-4B ties its embeddings, \(W_a \approx I\), which is exactly why this model was chosen:
$$e \;=\; \mathrm{RMSNorm}\big(\hat h_{\text{fin}}\big)\, W_a, \qquad W_a = \arg\min_W \big\lVert W_U W - W_E \big\rVert_F^2 \;\approx\; I$$Third, a ridge-regression matrix \(J^{\text{in}}\), fit on WikiText prompts, stands in for the sensory layers, mapping the embedding straight to a workspace-entry state:
$$h_{\text{in}} \;=\; \mu_{\text{ent}} \;+\; J^{\text{in}}\,\big(e - \mu_e\big)$$(The \(\mu\)'s are calibration means over WikiText; \(J^{\text{in}}\) has large gain, so \(h_{\text{in}}\) is rescaled to the typical entry-state norm to stay on the workspace manifold.) Then the state runs through the workspace only:
$$h_{28} \;\leftarrow\; \mathrm{Layers}_{6\ldots28}\big(h_{\text{in}}\big)$$and the loop repeats for \(K\) latent steps. The KV cache grows only in the workspace layers; the latent positions literally do not exist for the sensory and motor layers. After \(K\) steps, the cue "The answer is" is fed through the full stack and the answer is decoded normally. (Sanity check: with \(K=0\) the manual single-position runner reproduces model.generate token-for-token.)
On GSM8K (200 problems, greedy decoding, 2-shot): answering directly with \(K=0\) latent steps scores 25%. Turning the latent loop on lifts this to ≈31%, but the curve is flat in \(K\): 32.0% at \(K=8\), 30.5% at 16, 32.0% at 32, 31.5% at 64. More workspace-only thinking steps buy nothing. Text CoT at a matched token budget scales the way you'd want thinking to scale: 17% at 8 tokens, 23% at 16, 33% at 32, 58% at 64, and 89% uncapped (the natural CoT length here is ≈69 tokens). The crossover is around 24 tokens: below that budget, latent workspace thinking wins; above it, text wins and keeps climbing.
So the mechanics work: the workspace can be driven in a closed loop through purely linear bridges without leaving its manifold, and doing so delivers a real one-time lift over answering directly. What it doesn't do, training-free, is compound: the loop seems to hand the model a single "gist" of thought rather than accumulating computation step over step. The obvious suspects are the bridges themselves (linear, calibrated on WikiText, nothing task-specific); training them, or re-entering at the KV level instead of the residual stream, is where I'd push next.