Skip to content

Advancing Egocentric Video Dialogue: A Contextual Reasoning Approach with New Benchmark Dataset

Sep 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · 0 citations · 52 references

TL;DR

This work proposes a novel task named Egocentric Video Dialogue, which refers to answering questions by considering both the video content and dialogue history, and designs a new dataset, EgoDialogue, which constructs continuous dialogues based on video question answering datasets.

Abstract

Egocentric video reasoning capability is essential for advancing the development of first-person wearable devices. However, most existing egocentric video understanding tasks are limited to single-instance text input, overlooking the model's capability to reason about context throughout the entire duration of the video. To bridge this gap, we propose a novel task named Egocentric Video Dialogue, which refers to answering questions by considering both the video content and dialogue history. To support this task, we design a new dataset, EgoDialogue, which constructs continuous dialogues based on video question answering datasets. Egocentric videos and dialogue histories present two challenges, requiring the model to understand core interactions in visual data and the dialogue history in text, respectively. To address these issues, we introduce an EgoReason framework, consisting of an interaction exploration module and a dialogue reasoning module. The former module first reorganizes patches that correspond to the same spatial position across different frames and then sorts these blocks according to their frame order. In this way, it models spatiotemporal dynamics to help explore interactions. The latter module improves the understanding of dialogues by gradually integrating the questions and answers from each dialogue round into the fusion layer. In the experiments, our method shows superior performance compared to several VideoQA models and vision language models, achieving state-of-the-art results in prediction accuracy and various machine translation metrics.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Companion-style QA Assistance in Ego-Vision

Experiments show that MyBuddy significantly enhances the performance of foundation models on BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video, and generalize to other streaming and common video QA benchmarks, demonstrating the applicability and effectiveness of the approac...

Hangyu Qin, Jun-Bin Xiao, Sheng Zhang et al. · 0 citations
#small language model Review Aug 2026

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill...

M. Zamani, Fatemeh Ziaeetabar · 0 citations
Open access 2026

PAIR: A Pilot Dataset for Dual Perspective-based Video-Grounded Dialogue and Reconciliation

Collaborative dialogue in multi-agent settings often requires interlocutors to integrate partially overlapping perceptual information in order to construct a shared representation of a dynamic environment. We introduce PAIR, a pilot conversational corpus designed to examine how humans coordinate under systematic percep...

Lewis Watson, Carl Strathearn, Kenny Mitchell et al. · 0 citations
Preprint Sep 2026

Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding

Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and t...

Wei Chen, Xuan-Yu Zheng, Yan-Cheng Long et al. · 0 citations
Review Open access Sep 2026

A Survey of Multi-Model Collaboration in Video Understanding

The rapid development of multimodal foundation models has shifted video understanding from perception-centered recognition toward more general semantic interpretation, reasoning, and decision-making over dynamic visual content. As video understanding tasks increasingly require fine-grained perception, long-range tempor...

Yi Chen, Jian-Wei Zhang, Lei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.