We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain speci...
Ming-Fei Gao, Rui Tian, Hai-Ming Gang et al.· 0 citations
OmegaUse-OfficeVal is introduced, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding, and code-based verifiers from fine-grained rubrics are developed to support stable evaluation.
Jing-Bo Zhou, Yu-Sai Zhao, Qi Bao et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.