This work introduces a transparent SAE-feature steering pipeline that first applies a six-condition reliability filter, then ranks sparse features through an unweighted Borda consensus over three complementary statistics: F -test, KSG mutual information, and Cohen's d .
Steering Vector Dissection is proposed, a framework to explicitly isolate individual and semantically consistent features from these composite directions that tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects.
Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi et al.· 2 citations
This work introduces Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning and evaluates TFTVs on large language model behavioral control tasks and shows that they consistently amplify, suppress, and compose target behaviors while preserving general know...
G. Perin, Lucas Boscaini, André Araújo et al.· 0 citations
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-valu...
This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we investigate whether additive, training-free composition of attribute steering vectors can preserve the intended steering effect of each attribut...
Hyunku Kang, Daniil Gurgurov, Tanja Baeumel et al.· 0 citations
Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties...
Irene Tallini, Lorenzo Basile, Valentino Maiorca et al.· 0 citations
Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector across all activations, limiting expressivity and ignoring t...
Laziz U. Abdullaev, M. Pham, B. Do et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.