Skip to content

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

This work proposes RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation.

Abstract

Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs'feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.

View source

Similar papers

#machine learning Preprint Oct 2026

LocusRL: Diagnosing LLM Reward and Policy Interventions in Competitive Games

Large language models can intervene in reinforcement learning through both reward design and action selection, yet aggregate performance offers an incomplete account of what these interventions actually do. Similar returns can conceal different learning mechanisms, while plausible rewards can induce undesirable behavio...

Cheng-Yu Luan, Bo Xin, Song-Yan Guo et al. · 0 citations
#machine learning Preprint Sep 2026

Cliff: Learning Process Rewards from the First Mistake

Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout, is proposed and established as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.

Pei-Xuan Han, Runnan Wang, Ketan Ramaneti et al. · 1 citation
Review Open access Sep 2026

Foundation-Model-Assisted Reward Design for Reinforcement Learning: A Review of Reward Program Synthesis, Multimodal Feedback, and Trustworthiness

Reward functions determine what reinforcement learning agents ultimately optimize, yet reward design for complex tasks has traditionally relied on extensive domain expertise and iterative engineering. Recent large language models and vision–language foundation models have introduced new mechanisms for interpreting task...

Wei Zhu, Jin-Yin Bai, Rui Tang et al. · 0 citations
Preprint Sep 2026

HiRE: Hindsight Reward Editing for Policy Finetuning

Hindsight Reward Editing (HiRE) bridges the broad knowledge of foundation representation models with physical control awareness, by contrasting successful and failed trajectories in hindsight, and consistently outperforms other reward recipes by delivering dense, control-aware feedback that prevents value function coll...

Haoyi Niu, Zhen Han, Yu-Feng Ji et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's...

Zeng-Huang Fu, Zhao-Yang Li, Qiu-Yuan Ai et al. · 0 citations
#artificial intelligence Preprint Oct 2026

EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans

Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses...

Xuan-Cheng Li, Bei-Ning Wang, Hai-Tao Li et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.