Skip to content

Disentangling Steering Vectors

Sep 2026 · 2 citations · 37 references
Computer Science

TL;DR

Steering Vector Dissection is proposed, a framework to explicitly isolate individual and semantically consistent features from these composite directions that tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects.

Abstract

Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional steering vectors used to intervene in LLMs'activations, such as those derived from the difference-in-means method, tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects. Our core objective is to disentangle this composite direction into its constituent concepts. To this end, we propose Steering Vector Dissection, a framework to explicitly isolate individual and semantically consistent features from these composite directions. Specifically, we pair positive and negative activations and take their differences to generate a set of instance-level steering vectors, and train a dedicated Sparse Autoencoder (SAE) directly on them. Quantitative evaluations across two datasets, two models, and two intervention depths show that our method yields a set of semantically consistent basis vectors whose steering effects are mutually distinguishable. Furthermore, we show that this disentanglement enables precise control over model behaviors.

View source

Similar papers

#natural language process... Preprint Sep 2026

Compositional Multilingual and Behavioral Attribute Steering

This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we investigate whether additive, training-free composition of attribute steering vectors can preserve the intended steering effect of each attribut...

Hyunku Kang, Daniil Gurgurov, Tanja Baeumel et al. · 0 citations
#machine learning Preprint Sep 2026

LocUS: Head Selection and Subspace Projection for Targeted Activation Steering

Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties...

Irene Tallini, Lorenzo Basile, Valentino Maiorca et al. · 0 citations
#machine learning Preprint Oct 2026

Kernelized Activation Steering

Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector across all activations, limiting expressivity and ignoring t...

Laziz U. Abdullaev, M. Pham, B. Do et al. · 0 citations
#machine learning Preprint Sep 2026

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-valu...

Payel Bhattacharjee, Ravi Tandon · 0 citations
#artificial intelligence Preprint Sep 2026

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it u...

Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Signatures of Steerability in Activation Space of Language Models

Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properti...

Prajjwal Bhattarai, Tuka Alhanai · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.