Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning

Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segm...

Xin-Chen Du, Zheng-Ze Zhou, Wen-Hui Zhu et al. · 0 citations
#natural language process... Preprint Sep 2026

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth...

De Xu, B. Li, Bang Lin et al. · 35 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.