← back to the mech-interp place
An early-stage experiment: activation steering where you never have to sweep layers to find the injection point.
The standard recipe for steering a model with a vector (add \(v\) to the residual stream and the behavior shifts) hides a grid search. The vector only works when added at the right layer, at the right strength, so every new model and every new behavior starts with a sweep: inject at layer 5, evaluate, layer 6, evaluate, and so on until something sticks. The result is a single hand-picked layer that is brittle, model-specific, and expensive to find.
The idea here is to make the placement part of the mechanism instead of a hyperparameter. An attention mechanism scores each layer's residual state against the steering vector and distributes the vector across the whole stack in one pass:
$$\alpha \;=\; \mathrm{softmax}_{\,\ell}\big(\,\mathrm{score}(h_\ell,\, v)\,\big), \qquad h_\ell \;\leftarrow\; h_\ell \,+\, \alpha_\ell\, \beta\, v$$The model itself decides where the vector lands, and the injection can spread over several adjacent layers instead of concentrating in one hand-picked spot, which no sweep can express. The sweep over layers disappears; only the overall strength \(\beta\) remains.
Status: early-stage. The repo currently holds a Qwen2.5-1.5B harness and the idea. The interesting questions are still open: does the attention pick the same layers a sweep would, does spreading beat concentrating, and do the learned weights transfer across behaviors within one model.