Skip to content

Steering Instruction Hierarchies at Inference Time

Jul 2026 · arXiv.org · Vol abs/2607.26228 · 2 citations · 37 references
Computer Science

TL;DR

V-Steer, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions, is introduced, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions.

Abstract

Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs from users or tools. Yet frontier LLMs often violate this hierarchy. We introduce V-Steer, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions. Using direct logit attribution on the first next token prediction, V-Steer identifies heads where lower priority spans dominate privileged ones, then boosts privileged spans and suppresses conflicting lower priority spans through in-place multiplicative edits to cached V tensors. Since the method acts only on cached values, it remains compatible with fused attention backends and adds only a one time prefill overhead. Across models from 7B to 70B, this attribution guided intervention raises primary constraint accuracy from under 18% up to 92% on controlled role conflict benchmarks, and on broader instruction hierarchy evaluations substantially outperforms prompt only baselines while matching or exceeding SoTA training based methods on 3 of 4 scales of LLMs, with negligible decoding-speed overhead. The code is available at https://github.com/cindy2000sh/v-steer.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

It is shown that user-preferring conflict resolution can coexist with a readable internal arbitration signal, and successful intervention depends on the geometry of the readout rather than probe accuracy alone, while directions selected mainly for pooled separability steer poorly.

Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems

This work argues that a general workload representation, a verifiable mutation space, and an implementation-independent evaluator are required to enable the AI-driven LLM inference system architecting loop, and presents the RoofLang domain-specific language (DSL) that provides these features.

Ziyue Yang, Yu-Ting Jiang, L. Qu et al. · 0 citations
Preprint Aug 2026

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

This work introduces Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.

Zining Huang, Haoran Que, Hongxia Zeng et al. · 2 citations
#machine learning Preprint Sep 2026

JET: Justification Evaluation in Transformer

The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference, and support local decision inference from existing models supports local decision inference from existing models.

Sheng-Hao Ding · 0 citations
#natural language process... Preprint Sep 2026

ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation

Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organize...

Hui-Fei Wang, Xin-Yi Huang, Yi-Heng Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation

This work presents vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization, and introduces IMPACT, an ACT-based policy with cached text representations and language-modulated visual features.

Khanh Duy Nguyen, Hoang M. Truong, A. T. Le · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.