Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin, and the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.
Ding-Yu Yao, Jun Zhou, Chenxu Yang et al.· arXiv.org· 10 citations
VIABench is introduced, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves, and proposes a rigorous benchmarking pipeline that supports both online (real-time) and offline settings.
Yunfeng Liu, Yuandong Yang, Jiarui Han et al.· arXiv.org· 0 citations
This work introduces VideoChat3, a fully open, efficient, and generalist video-centric MLLM, which surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.
Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters.