Skip to content

Author

Guanghua Yu

We have 6 of 20 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

D-Quant: Driftable Entropy Coding for KV Cache Quantization

The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is particularly attractive due to its effectiveness a...

Yi Su, Hong Liu, Guang-Hua Yu et al. · 0 citations
Jul 2026

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention

Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2...

Hong Liu, Yuan Cheng, Lin Niu et al. · 1 citation · ⚡1
Preprint Aug 2026

FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling

FOCUS is proposed, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling, which relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimiza...

Xiang-Long Yan, Hong Liu, Cheng-Zhu Bao et al. · 1 citation
Jul 2026

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

HuyuanOCR-1.5 ranks among the top-tier end-to-end OCR solutions on OmniDocBench v1.6 while achieving new performance milestones across these long-tail tasks, and proposes Agentic Data Flow, an agent-driven data construction system that transforms model weaknesses into executable data requirements and autonomously perfo...

Gengluo Li, Xingyu Wan, Shangpin Peng et al. · 7 citations · ⚡1
Jul 2026

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

CoSA is proposed, a two-stage training-free Sparse Attention under proxy-kernel CO-design, which couples a Kernel-Aware Proxy (KAP) with an Ordered-Skipping Kernel (OSK) and achieves a 4.93% attention speedup and reduces end-to-end Time-to-First-Token by 2.53% with negligible performance degradation.

Yufei Xue, Lin Niu, Hong Liu et al. · 0 citations
Jul 2026

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

DFly is proposed, a block-diffusion framework combining a hybrid target-conditioning backbone with a predecessor-conditioned autoregressive head, improving target-feature utilization and intra-block dependency modeling while keeping generation parallel, and DFly treats verification as a shared batch-level resource.

Hong Liu, Rui Cen, Jun-Han Shi et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.