Active 3D Gaussian reconstruction fundamentally relies on selecting informative next-best views under limited sensing budgets. Existing active 3DGS methods primarily plan viewpoints according to geometric information gain, treating object-induced hidden regions in the same manner as general unexplored space. Under tigh...
Hong-Bo Gao, Wei Zhang, Ze-Yu Ni et al.· 0 citations
MoDeVLA is proposed, the first rate-distortion driven efficient VLA model that performs token-wise depth allocation via Mixture-of-Depth Conditioning and integrates shallow visual-spatial with deep textual-logical features for action conditioning and introduces Effective-Edge Flow, an action-aligned attribution measure...
Weiying Xie, Qingcheng Zeng, Zihan Meng et al.· Proceedings of the 32nd ACM...· 0 citations
Vision-Language-Action (VLA) model plays a crucial role in embodied decision making. While practical deployment requires fast inference under limited onboard computation, a full forward pass through the vision-language model makes such deployment challenging. To address this issue, existing methods typically employ lig...
Weiying Xie, Qingcheng Zeng, Zihan Meng et al.· Proceedings of the 32nd ACM...· 0 citations
LoSA is a training-free sparse-attention method that fixes a retained-mass threshold of 99% rather than a sparsity ratio: it measures exact block attention masses at one early dense step, keeps, for each head and query block, the smallest key/value block set meeting the threshold, and reuses the frozen block indices fo...
En-Huai Liu, Yun-Ke Wang, Yu-Tong Wang et al.· 0 citations
The real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.
Siyu Xu, Yun-Ke Wang, Zi-Jian Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.