Model Diffing: CLIP vs. DINOv3
My main research right now. Diffing vision models CLIP and DINOv3 to understand how CLIP’s joint language objective costs it fine-grained visual features. Designed a novel tournament-style top-K algorithm for block-sparse SAEs, significantly reducing feature splitting. Experimented with block-sparse cross-coders, alongside Matryoshka and Archetypal SAEs, in order to find the best diffing recipe. With Amith Ananthram (Anthropic Fellow), in Prof. Kathleen McKeown’s lab; in preparation for ICLR.