Skip to content

Author

Qing-Peng Cai

We have 6 of 18 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

Generative Auto-Bidding in Large-Scale Auctions via Diffusion Completer-Aligner

Bid optimization strategy in auto-bidding is central to computational advertising, achieving notable commercial success by optimizing advertisers' bids within constraints. Recently, generative models have revolutionized auto-bidding by directly learning a policy from large-scale datasets. Among them, the diffuser is su...

Ye-Wen Li, Jingtong Gao, Peng Jiang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

Auto-bidding is central to computational advertising, where strategies must maximize advertisers'conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevail...

Ye-Wen Li, Peng Jiang, Yi-Tian Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs

Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when several outputs jointly cause one state transition. Consequently, the update unit differs fr...

Sheng-Tian Yang, Zi-Yun Xiong, Yu Li et al. · 1 citation
#artificial intelligence Preprint Sep 2026

AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

AgentBrew is proposed, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts, and demonstrates that fine-grained offline learning can recover useful supervision from raw trajectories that filterin...

Zhiyi Lyu, Ye-Wen Li, Longtao Zheng et al. · 2 citations
#machine learning Preprint Aug 2026

PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization

This work proposes PRQ-KMeans, which removes the global-mean component, refines centroids with Top-k similarity-weighted updates, and replaces full-codeword subtraction with a projection residual that removes each representation's selected-centroid component.

Yunxiao Luo, Siyuan Wang, Ben Chen et al. · 0 citations
Book Open access Aug 2026

Hierarchical Residual Policy Optimization for Generative Recommendations

Hierarchical Residual Policy Optimization (HRPO), a post-training framework that converts item-level outcomes into dense, token-aligned learning signals for conservative token-wise improvement, is proposed.

Kaifeng Guo, Yiming Yang, Jingtong Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.