Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may...
Shan-Wen Mao, Ming-Ming Li, Hao Zhang et al.· 0 citations
This work proposes Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization to address optimization conflicts among multiple objectives in real-world deployment scenarios.
Shang-Wen Mao, Hao Zhang, Guangtao Nie et al.· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.