Preprint
Aug 2026
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Capek 0.5 is presented, an embodied vision-language model built around an execution-centric capability taxonomy that improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.
Ying Chen, Weizhen Li, Zhe Hu et al.
· 0 citations