Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not charac...
Xinming Wang, Wei-Nong Wang, Hongming Yang et al.· 1 citation· ⚡1
As large multimodal models move from understanding content to operating on digital environments, mobile GUI has emerged as a challenging and consequential testbed for digital embodied intelligence. Mobile agents operate under three coupled constraints: precise perception of complex interfaces, scalable acquisition of h...
Hy Vision Team, Huawen Shen, Zhengyang Tang et al.· arXiv.org· 1 citation
HuyuanOCR-1.5 ranks among the top-tier end-to-end OCR solutions on OmniDocBench v1.6 while achieving new performance milestones across these long-tail tasks, and proposes Agentic Data Flow, an agent-driven data construction system that transforms model weaknesses into executable data requirements and autonomously perfo...