Skip to content

Author

Si-Rui Han

We have 4 of 17 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present O...

De-Zhi Li, Lu-Jun Li, Qi-Yuan Zhu et al. · 0 citations
#natural language process... Preprint Sep 2026

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.

Zhi-Wei Li, Lei Zhu, Hao Gu et al. · 0 citations
Preprint Jul 2026

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

Hierarchical Landmark Sparse Attention is proposed, a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling (LM) loss, enabling long-context LLMs that are both more efficient and more effective on general long-context tasks than their full-attention counterparts.

Xiang Hu, Xinyu Wei, Hao Gu et al. · 3 citations
Preprint Aug 2026

Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning

ReCo (Reward-Coordinated Compression), a step-wise framework in which a lightweight process-reward estimator scores each completed step and drives three components: reward-adaptive KV-cache compression that shrinks the retained cache harder at high-reward steps and less at low-reward ones, and a confidence-based early...

Qi-Yuan Zhu, De-Zhi Li, Pengyu Cheng et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.