RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
This work proposes RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass gen...