Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. T...
Jun-Yi Wu, Fan-Qing Kong, Le-Yang Chen et al.· 0 citations
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-trainin...
Kai-Cheng Yang, Kai-Sen Yang, Chun-Yu Liu et al.· 0 citations
A causal interpretability framework is introduced that reveals that the tokens encoding semantics at the input need not be those that maintain them during generation, but also provides a causal view of internal mechanisms in diffusion transformers.
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence...
Zizhong Ding, Junxian Li, Kai Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.