WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents
WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery, and WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout that outperform similarly sized open-source models and rival models with roughly ten times its parameter count.
Zongkai Liu, Hui Zhang, Liqiang Niu et al.
· 0 citations