Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundan...
Hao Wu, Yang Xiao, Yu-Song Sun et al.· 0 citations
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositiona...
Xiao-Wen Yang, Wei-Yi Xu, Wen Da et al.· 0 citations
Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy, is developed.
Jia-Yan Fu, Hang Xu, Yong Zhang et al.· 0 citations
Platform-scale recommender systems often span multiple business streams such as organic recommendation, advertising, and merchant services, where user behaviors form a continuous cross-stream trajectory. Maintaining separate ranking systems fragments user representations and increases engineering cost. We propose \text...
Yinqi Zhang, Pei-Yu Hu, Yuntian Tang et al.· 1 citation
Agentic recommender systems use large language models to maintain semantic memory and support evidence-aware recommendation. However, existing memory mechanisms often compress user and item information into coarse summaries and connect them with scalar collaborative links, making it difficult to preserve fine-grained p...
Pei-Yu Hu, Wei-Hai Lu, Si-Ying Gu et al.· 0 citations
PILOT is presented, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and...
Yang Xiao, Yu-Song Sun, Haoming Wu et al.· 2 citations
To further optimize the reasoning trajectory, HiLaR combines final recommendation feedback with layer-aware process rewards derived from the marginal target-likelihood gain of each state, and generally outperforms strong sequential, generative, and LLM-based recommendation baselines.
Pei-Yu Hu, Si-Ying Gu, Wei-Hai Lu et al.· arXiv.org· 1 citation
This work proposes Akashic, a low-overhead memory system built around MemAttention, which organizes context into bounded chunks and models semantic relationships across chunks, preserving cross-chunk evidence without repeatedly rewriting the full history.
Yang Liu, ZhaoKai Luo, Huayi Jin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.