Recent advancements in Large Language Models (LLMs) have inspired their application in sequential recommendation systems, often through supervised fine-tuning (SFT). However, conventional SFT methods often struggle to capture nuanced comparative relationships between items. While recent approaches utilize Direct Preference Optimization (DPO), they remain constrained by challenges in modeling customized preferences, including capturing fine-grained user preferences and being susceptible to semantic hallucination and popularity bias. To overcome these challenges, we propose RosePO, a framework to refine LLM-based recommendation through pairwise preference optimization with personalized smoothing. We illustrate the concept with three critical preference examples pertinent to LLM-based recommendation. Specifically, we design rejected sampling strategies tailored for each customized preference. To ensure robustness against uncertain labels present in automatically constructed preference data, we incorporate a personalized smoothing factor predicted by a user oracle into the optimization objective. Empirical evaluation on three real-world datasets demonstrates the effectiveness of our method, showcasing not only promising recommendation performance but also mitigation of semantic hallucination and popularity bias. We hope this work paves a way to build helpful and harmless LLM-based recommendation service in the future.
Jiayi Liao, Xiangnan He, Ruobing Xie et al.· ACM Transactions on Informat...· 0 citations
This work proposes ARMOR (Anchor Rollout and Mixed Optimization for RL), a framework that shifts the paradigm from passive penalty to active sample stabilization, enabling sustained performance improvements over extended training horizons.
Kexin Huang, Junkang Wu, Jinda Lu et al.· 0 citations
Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity is proposed, which demonstrates that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base model.
Jiawei Feng, Jiancan Wu, Xingyu Zhu et al.· 1 citation