We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T...
Tencent Hunyuan Team, Ao Liu, Bo Zhou et al.· 0 citations
The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is particularly attractive due to its effectiveness a...
HPC-Ops Top-K is presented, a sample-guided exact selector for ragged sparse-attention score rows that outperforms the fastest verified external exact baseline and outperforms the fastest verified external exact baseline on indexer scores from Hy4-Preview.
Si-Ran Liu, Ya-Long Xue, Theo Tang et al.· 0 citations
Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2...
Hong Liu, Yuan Cheng, Lin Niu et al.· arXiv.org· 1 citation· ⚡1
FOCUS is proposed, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling, which relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimiza...
Xiang-Long Yan, Hong Liu, Cheng-Zhu Bao et al.· 1 citation
CoSA is proposed, a two-stage training-free Sparse Attention under proxy-kernel CO-design, which couples a Kernel-Aware Proxy (KAP) with an Ordered-Skipping Kernel (OSK) and achieves a 4.93% attention speedup and reduces end-to-end Time-to-First-Token by 2.53% with negligible performance degradation.
Yufei Xue, Lin Niu, Hong Liu et al.· arXiv.org· 0 citations
DFly is proposed, a block-diffusion framework combining a hybrid target-conditioning backbone with a predecessor-conditioned autoregressive head, improving target-feature utilization and intra-block dependency modeling while keeping generation parallel, and DFly treats verification as a shared batch-level resource.
Hong Liu, Rui Cen, Jun-Han Shi et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.