Sep 2026· ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP)· 0 citations· 52 references
TL;DR
This work proposes a novel task named Egocentric Video Dialogue, which refers to answering questions by considering both the video content and dialogue history, and designs a new dataset, EgoDialogue, which constructs continuous dialogues based on video question answering datasets.
Abstract
Egocentric video reasoning capability is essential for advancing the development of first-person wearable devices. However, most existing egocentric video understanding tasks are limited to single-instance text input, overlooking the model's capability to reason about context throughout the entire duration of the video. To bridge this gap, we propose a novel task named Egocentric Video Dialogue, which refers to answering questions by considering both the video content and dialogue history. To support this task, we design a new dataset, EgoDialogue, which constructs continuous dialogues based on video question answering datasets. Egocentric videos and dialogue histories present two challenges, requiring the model to understand core interactions in visual data and the dialogue history in text, respectively. To address these issues, we introduce an EgoReason framework, consisting of an interaction exploration module and a dialogue reasoning module. The former module first reorganizes patches that correspond to the same spatial position across different frames and then sorts these blocks according to their frame order. In this way, it models spatiotemporal dynamics to help explore interactions. The latter module improves the understanding of dialogues by gradually integrating the questions and answers from each dialogue round into the fusion layer. In the experiments, our method shows superior performance compared to several VideoQA models and vision language models, achieving state-of-the-art results in prediction accuracy and various machine translation metrics.
Experiments show that MyBuddy significantly enhances the performance of foundation models on BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video, and generalize to other streaming and common video QA benchmarks, demonstrating the applicability and effectiveness of the approac...
Hangyu Qin, Jun-Bin Xiao, Sheng Zhang et al.· 0 citations
This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill...
Collaborative dialogue in multi-agent settings often requires interlocutors to integrate partially overlapping perceptual information in order to construct a shared representation of a dynamic environment. We introduce PAIR, a pilot conversational corpus designed to examine how humans coordinate under systematic percep...
Lewis Watson, Carl Strathearn, Kenny Mitchell et al.· International Conference on...· 0 citations
Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and t...
Wei Chen, Xuan-Yu Zheng, Yan-Cheng Long et al.· 0 citations
The rapid development of multimodal foundation models has shifted video understanding from perception-centered recognition toward more general semantic interpretation, reasoning, and decision-making over dynamic visual content. As video understanding tasks increasingly require fine-grained perception, long-range tempor...
Yi Chen, Jian-Wei Zhang, Lei Zhang et al.· International Conference on...· 0 citations
A cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios, and two official Codabench tracks.
Yu-Qian Fu, Tianwen Qian, Yanjun Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.