Skip to content

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

2026 · arXiv.org · Vol abs/2607.19364 · 0 citations · 31 references
Computer Science

TL;DR

This work introduces a transparent SAE-feature steering pipeline that first applies a six-condition reliability filter, then ranks sparse features through an unweighted Borda consensus over three complementary statistics: F -test, KSG mutual information, and Cohen's d .

View source

Similar papers

#machine learning Preprint Sep 2026

Disentangling Steering Vectors

Steering Vector Dissection is proposed, a framework to explicitly isolate individual and semantically consistent features from these composite directions that tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects.

Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Training-Free Task Vectors for LLM Behavioral Control

This work introduces Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning and evaluates TFTVs on large language model behavioral control tasks and shows that they consistently amplify, suppress, and compose target behaviors while preserving general know...

G. Perin, Lucas Boscaini, André Araújo et al. · 0 citations
#machine learning Preprint Sep 2026

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-valu...

Payel Bhattacharjee, Ravi Tandon · 0 citations
#natural language process... Preprint Sep 2026

Compositional Multilingual and Behavioral Attribute Steering

This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we investigate whether additive, training-free composition of attribute steering vectors can preserve the intended steering effect of each attribut...

Hyunku Kang, Daniil Gurgurov, Tanja Baeumel et al. · 0 citations
#machine learning Preprint Sep 2026

LocUS: Head Selection and Subspace Projection for Targeted Activation Steering

Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties...

Irene Tallini, Lorenzo Basile, Valentino Maiorca et al. · 0 citations
#machine learning Preprint Oct 2026

Kernelized Activation Steering

Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector across all activations, limiting expressivity and ignoring t...

Laziz U. Abdullaev, M. Pham, B. Do et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.