Skip to content

Author

Linlin Yang

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

GleanVID: Complementary Token Selection for Efficient Video Large Language Models

Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and conseque...

Shuo Yang, Chang-Bai Li, Rui Tang et al. · 0 citations
#machine learning Preprint Sep 2026

GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents

Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its r...

Hao-Dong Zhu, Yang-Yang Ren, Chang-Bai Li et al. · 0 citations
#machine learning Preprint Sep 2026

Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning

Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence be...

Yang-Yang Ren, Hao-Dong Zhu, Lin-Lin Yang et al. · 0 citations
Preprint Aug 2026

Contextual Information Policy Optimization for Search Agents

Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant evidence but also on using it t...

Xingyu Guo, Wei Chen, Lin-Lin Yang et al. · 0 citations
Jul 2026

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

A Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction, and consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt se...

Hao-Dong Zhu, Yang-Yang Ren, Yanjing Li et al. · 2 citations · ⚡2

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.