This work proposes RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation.
Abstract
Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs'feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.
Large language models can intervene in reinforcement learning through both reward design and action selection, yet aggregate performance offers an incomplete account of what these interventions actually do. Similar returns can conceal different learning mechanisms, while plausible rewards can induce undesirable behavio...
Cheng-Yu Luan, Bo Xin, Song-Yan Guo et al.· 0 citations
Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout, is proposed and established as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.
Pei-Xuan Han, Runnan Wang, Ketan Ramaneti et al.· 1 citation
Reward functions determine what reinforcement learning agents ultimately optimize, yet reward design for complex tasks has traditionally relied on extensive domain expertise and iterative engineering. Recent large language models and vision–language foundation models have introduced new mechanisms for interpreting task...
Hindsight Reward Editing (HiRE) bridges the broad knowledge of foundation representation models with physical control awareness, by contrasting successful and failed trajectories in hindsight, and consistently outperforms other reward recipes by delivering dense, control-aware feedback that prevents value function coll...
Haoyi Niu, Zhen Han, Yu-Feng Ji et al.· 0 citations
Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's...
Zeng-Huang Fu, Zhao-Yang Li, Qiu-Yuan Ai et al.· 0 citations
Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses...
Xuan-Cheng Li, Bei-Ning Wang, Hai-Tao Li et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.