Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. T...
Jun-Yi Wu, Fan-Qing Kong, Le-Yang Chen et al.· 0 citations
LT-OPD, a training framework for extreme visual-token reduction, is proposed and it is shown that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
Junxian Li, Rui-Xuan Yang, Tian-Ao Zhang et al.· 0 citations
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-trainin...
Kai-Cheng Yang, Kai-Sen Yang, Chun-Yu Liu et al.· 0 citations
SPHQuant is a rotation-free spherical weight-only quantization framework for VLMs that isolates outlier magnitude into the radius while keeping directions bounded and statistically regular and designs a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radia...
Ke-Wei Zhang, Zheng Chen, Hao-Tong Qin et al.· 0 citations
FOCUS is proposed, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling, which relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimiza...
Xiang-Long Yan, Hong Liu, Cheng-Zhu Bao et al.· 1 citation
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence...
Zizhong Ding, Junxian Li, Kai Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.