Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single direction in activation space often suffices for steering. However, many concepts are not binary: Animals and Countries contain many subcate...
Divya Appapogu, Freya Behrens, Yonatan Belinkov et al.· 0 citations
The history and current state of interpretability taxonomized according to the types of causal units utilized, as well as methods used to search over mediators are described.
Aaron Mueller, Jannik Brinkmann, Millicent Li et al.· 0 citations
Three lines of evidence are shown that provide convergent evidence in support of the interlingua hypothesis, which holds that language models translate by reading a source sentence into a latent feature space, and generate a target sentence by reading from the latent feature space.
Jacob Brinton, Jannik Brinkmann, Mark Crovella et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.