Embodied navigation requires agents to ground instructions or object goals in spatial observations and translate plans into successful execution. As multimodal large language models (MLLMs) become increasingly capable, they offer stronger support for navigation without task-specific training; however, improved semantic...
Yang Chen, Li-Rong Che, Zhen-Yu Huang et al.· 4 citations· ⚡1
CoBench is introduced, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks and shows that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types.
Yang Chen, Ye-Xin Xie, Li-Rong Che et al.· 0 citations
VERA (Visual Evidence-Retaining strategy for long-horizon Agents), a training-free context manager built on deterministic rendering with no exposed memory operations, supporting a modality-preserving view of long-horizon context management.
Jiang-Feng Su, Cong Pang, Jiawei Hong et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.