Skip to content

Author

Yulun Zhang

We have 6 of 15 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

VASC: Value-Aware Sparse Attention with Cross-Layer Memory for Efficient 3D Reconstruction

Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. T...

Jun-Yi Wu, Fan-Qing Kong, Le-Yang Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization

Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-trainin...

Kai-Cheng Yang, Kai-Sen Yang, Chun-Yu Liu et al. · 0 citations
#small language model Preprint Sep 2026

SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models

SPHQuant is a rotation-free spherical weight-only quantization framework for VLMs that isolates outlier magnitude into the radius while keeping directions bounded and statistically regular and designs a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radia...

Ke-Wei Zhang, Zheng Chen, Hao-Tong Qin et al. · 0 citations
Preprint Aug 2026

FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling

FOCUS is proposed, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling, which relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimiza...

Xiang-Long Yan, Hong Liu, Cheng-Zhu Bao et al. · 1 citation
Preprint Aug 2026

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence...

Zizhong Ding, Junxian Li, Kai Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.