Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VL...
Junghyun Kim, Ngseo Kim, Chung-Woo Lee et al.· 0 citations
Geometry-Change VLA (GC-VLA), which learns to predict multiview future-current geometry-change tokens from current observations and applies Geometry-Conditioned Residual Flow (GCRF), using a binary intervention router and a single bounded residual velocity policy learned from closed-loop feedback.
Jinu Pahk, Jesoon Kang, T. Park et al.· 0 citations
Scene-Q is presented, a confidence-aware coarse-to-fine querying framework that normalizes encoder scores with temperature scaling and selectively invokes a reasoning VLM only for low-confidence cases, and improves open-vocabulary 3D instance segmentation on ScanNet200 and natural-language 3D instance retrieval on real...
Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which require reasoning over semantic perturbations of objects, attributes, relations, and their interactions. However, our controlled analysis reveals that existing compositionality-aware VLMs exhibit element-specif...
Seongjun Jeong, Minjoon Jung, Woo-Suk Choi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.