Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental qu...