Skip to content

Author

Yongheng Deng

We have 5 of 30 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents

Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While effective in control...

Dong-Hua Cai, Yong-Heng Deng, Yi-Fei Wang et al. · 0 citations
Book Open access Aug 2026

Accelerating Graph-Based RAG Retrieval via Locality-Aware Device-Cloud Collaboration

This paper proposes Lever, a locality-aware collaborative retrieval framework that exploits query locality to accelerate graph-based RAG retrieval, and identifies and empirically validate a previously underexplored property of RAG workloads: strong per-user query locality.

Yong-Heng Deng, Tianyuan Jiang, Zhen-Ya Ma et al. · 0 citations
Book Open access Aug 2026

Personalizing Large Language Models with User Profile Memory

Large language models (LLMs) are increasingly used in personalized applications, where responses must align with individual user preferences, histories, and profiles. A common approach is to inject user information into the prompt at inference time. However, existing methods typically rely on flat profile representatio...

Yang-Xu Liao, Yong-Heng Deng, Tianyuan Jiang et al. · 0 citations
Book Open access Aug 2026

Accelerating Graph-Based RAG Retrieval via Locality-Aware Device-Cloud Collaboration

Retrieval-Augmented Generation (RAG) grounds large language models in external knowledge and has become a key technique for knowledge-intensive tasks. As knowledge bases continue to scale, however, the retrieval stage increasingly dominates end-to-end latency, limiting the responsiveness of RAG systems. In this paper,...

Yongheng Deng, Tianyuan Jiang, Zhenya Ma et al. · 0 citations
Book Open access Aug 2026

Personalizing Large Language Models with User Profile Memory

PersMem is proposed, a user profile memory framework for LLM personalization that addresses three key questions: what user information to store, how to organize it, and how to use it effectively during generation that consistently improves personalization effectiveness while reducing prompt length.

Yang-Xu Liao, Yongheng Deng, Tianyuan Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.