Skip to content
Book Open access

A Geometric Information Bottleneck for Activation Steering

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 30 references

TL;DR

This work proposes a disentanglement-based intervention framework, termed IB-ACT, that identifies both where and how to intervene by exploiting the layer-wise geometry of representation spaces to isolate behavior-relevant information without disturbing other dimensions, and introduces a layer-selection mechanism determined prior to intervention.

Abstract

Activation-based steering methods for large language models often induce broad, entangled changes in model behavior, inadvertently altering capabilities unrelated to the intended behavior, which limits their reliability for fine-grained behavioral control. We address this limitation by reframing behavioral intervention through a geometric information bottleneck (IB) perspective, in which effective steering corresponds to selectively modifying task-relevant information while preserving the geometric structure of orthogonal representational subspaces. Building on this view, we propose a disentanglement-based intervention framework, termed IB-ACT, that identifies both where and how to intervene by exploiting the layer-wise geometry of representation spaces to isolate behavior-relevant information without disturbing other dimensions. Our method introduces a layer-selection mechanism determined prior to intervention, rather than relying on post hoc sparsity or regularization losses, and applies geometrically constrained transformations that target behavior-relevant subspaces in activation space while preserving orthogonal structure. We provide theoretical justification showing that interventions at these layers reduce unintended information leakage under an IB-style objective. Empirically, we evaluate IB-ACT on toxicity control and hallucination reduction in large language models and demonstrate consistent improvements over recent baselines, while analyzing the spectral structure of behavior-relevant representations for jailbreak mitigation. Overall, our findings suggest that selectively intervening at structurally appropriate layers is critical for controllable and disentangled behavioral steering in large language models.

Read PDF

Similar papers

Preprint Aug 2026

Rewriting or Reweighting? A Geometric Account in Language Models

Behavior manifold analysis is introduced, which isolates behavior-specific geometry by selecting sparse behavior-associated coordinates and lifting them into low-dimensional local charts, and provides a unified framework for understanding the mechanistic distinction between the two objectives.

Jun-Tong Wang, Shengkun Yang, Xiyuan Wang et al. · 0 citations
#machine learning Preprint Sep 2026

LocUS: Head Selection and Subspace Projection for Targeted Activation Steering

Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties...

Irene Tallini, Lorenzo Basile, Valentino Maiorca et al. · 0 citations
#machine learning Preprint Sep 2026

Disentangling Steering Vectors

Steering Vector Dissection is proposed, a framework to explicitly isolate individual and semantically consistent features from these composite directions that tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects.

Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it u...

Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari et al. · 0 citations
#natural language process... Preprint Sep 2026

GAPS: Dimension-Level Gates for Conditional Activation Steering

Activation steering suppresses undesired behaviors in language models by adding a steering vector to the hidden state during generation. Recent conditional methods such as CAST and DSAS improve the behavior-capability trade-off by deciding when to intervene, but once active, they apply the full dense vector to all hidd...

Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.