On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from the student's state....
Ang-Qing Jiang, Gao-Ming Zhang, Chao-Qun Zhang et al.· 0 citations
Dense retrieval has become a cornerstone of modern local-lifestyle e-commerce search by encoding queries and items into semantic embedding spaces. While recent advancements have transitioned from BERT-based embedding models to Large Language Models (LLMs), most approaches still treat LLMs as static text encoders, negle...
Ang-Qing Jiang, Gao-Ming Zhang, Jian-Chun Song et al.· 2 citations
A Hierarchical Semantic Alignment module to align query's latent space with item's quantization path and synchronize multi-granular semantics, and a personalized GR framework that models user behavior by synergizing discrete SIDs for structural guidance and continuous representations for fine-grained semantic refinemen...
Gao-Ming Zhang, Ang-Qing Jiang, Jian-Chun Song et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.