Skip to content

Author

Kai-Xin Ma

We have 2 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

MintAct: A Unified Visual Agent for Digital Environments

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain speci...

Ming-Fei Gao, Rui Tian, Hai-Ming Gang et al. · 0 citations
Preprint Jul 2026

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

Evaluating 12 state-of-the-art models, from 4B open-weight to frontier proprietary systems, shows that current models still lack robust visual tool-calling capability: even the best model achieves below 50% success rate, suggesting fundamentally different research directions for improving models at different capability...

Kaixin Ma, Di Feng, Alexander Metz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.