The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the described target is absent. We present a simple staged solution. Qwen3-ASR first converts speech into text. Several video mask tracks are then produced with complementary grounding and segmentation models. Instead of trusting a single prediction, we select the track that has the highest average mask agreement with the other candidates. A small set of explicit direction, count, and plural rules corrects queries that require more than ordinary single-object tracking. Finally, a video-level classifier combines visual, audio-visual, and within-video query scores to decide whether any target is present. The submitted system obtains 0.5952 J&F, 0.7931 no-target accuracy, 0.9205 target accuracy, and a final score of 0.769589. The challenge organizers notified our team that this result ranked first in the track.
Yiwen Ren, Jianing Liu, Yingxin Wang et al.· 0 citations
A five-dimensional analytical framework is introduced that clarifies how EI transforms isolated search trajectories into cumulative scientific insight, and identifies critical bottlenecks regarding evaluation, process traceability, and shared infrastructure.
Chao Wang, Lingling Li, Fang Liu et al.· 0 citations