Preprint
Jul 2026
Implicit Reasoning Steering via Concept Chaining
The results show that indirect, natural-looking text can systematically steer model predictions while remaining substantially less inferable than direct paraphrases, which shows that reasoning brittleness is not merely an evaluation artifact: it creates a practical channel through which latent biases can be amplified by ordinary-looking text to covertly redirect model decisions.
Xiao Ye, Sanika Chavan, Yuxi Huang et al.
· 0 citations