Bid optimization strategy in auto-bidding is central to computational advertising, achieving notable commercial success by optimizing advertisers' bids within constraints. Recently, generative models have revolutionized auto-bidding by directly learning a policy from large-scale datasets. Among them, the diffuser is su...
Ye-Wen Li, Jingtong Gao, Peng Jiang et al.· Proceedings of the 32nd ACM...· 0 citations
Auto-bidding is central to computational advertising, where strategies must maximize advertisers'conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevail...
Ye-Wen Li, Peng Jiang, Yi-Tian Li et al.· 0 citations
Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when several outputs jointly cause one state transition. Consequently, the update unit differs fr...
Sheng-Tian Yang, Zi-Yun Xiong, Yu Li et al.· 1 citation
AgentBrew is proposed, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts, and demonstrates that fine-grained offline learning can recover useful supervision from raw trajectories that filterin...
Zhiyi Lyu, Ye-Wen Li, Longtao Zheng et al.· 2 citations
This work proposes PRQ-KMeans, which removes the global-mean component, refines centroids with Top-k similarity-weighted updates, and replaces full-codeword subtraction with a projection residual that removes each representation's selected-centroid component.
Yunxiao Luo, Siyuan Wang, Ben Chen et al.· 0 citations
Hierarchical Residual Policy Optimization (HRPO), a post-training framework that converts item-level outcomes into dense, token-aligned learning signals for conservative token-wise improvement, is proposed.
Kaifeng Guo, Yiming Yang, Jingtong Gao et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.