Skip to content

Training-Free Task Vectors for LLM Behavioral Control

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

This work introduces Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning and evaluates TFTVs on large language model behavioral control tasks and shows that they consistently amplify, suppress, and compose target behaviors while preserving general knowledge and problem-solving skills.

Abstract

Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. To address this limitation, we introduce Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning. Our method maps activation steering vectors to rank-one weight-space edits using only forward-pass statistics, while satisfying arithmetic properties that directly support learning via addition, forgetting via subtraction, and the composition of multiple edits. Empirically, we evaluate TFTVs on large language model behavioral control tasks and show that they consistently amplify, suppress, and compose target behaviors while preserving general knowledge and problem-solving skills. We also validate our method against other editing and steering baselines, experimentally demonstrating that TFTVs achieve stronger trait control with better or competitive utility preservation. We hope our work opens new directions for the community in post-training model editing and broader training-free model control. Code is available on the project website: tftv-llm.github.io.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the o...

Naveen Vakada, Ming-Yuan Li, Shao-Xiong Ji · 0 citations
#artificial intelligence Preprint Sep 2026

Learn Here, Move Less Elsewhere: Input-Conditioned Plasticity from Retained-Domain Activation Atlases

ATLAS is introduced, which turns retained-domain representations into an input-dependent rule for task adaptation, and provides a practical mechanism for acquiring specialized skills while maintaining continuity in existing responses.

Jiang-Tao Lin, Bang-Yang Wei, Yi-Hang Ding et al. · 0 citations
#machine learning Preprint Sep 2026

Understanding and Exploiting Anisotropy in Post-Training

LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large a...

Samyak Jha, Harshvardhan Saini, Yi-Zhen Liao et al. · 0 citations
Preprint Aug 2026

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

Evaluation-Conditioned Training (ECT), a post-training framework that uses natural language to condition each training sample on the fidelity of the feedback the authors provide and then elicits the desired behavior by conditioning the LLM on a high-fidelity monitor in deployment, is introduced.

Alec Harris, Kasey Corra, Archie Chaudhury et al. · 0 citations
#machine learning Preprint Sep 2026

Online Self-Weighted Fine-Tuning

Online Self-Weighted Fine-Tuning is proposed, a simple method that augments SFT with online, trajectory-level weighting and offers a favorable compute-performance trade-off as a practical approach for fine-tuning small-to-medium LLMs on binary-verifiable reasoning tasks with only 2 online rollouts.

Hai-Quan Wen, Yiwei He, Bei Peng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Behavioral Foundation Models for Quality Diversity

Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a la...

Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.