Experiments show that MyBuddy significantly enhances the performance of foundation models on BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video, and generalize to other streaming and common video QA benchmarks, demonstrating the applicability and effectiveness of the approach.
Abstract
AI companions are envisioned as always-on assistants that support users in daily life. With this regard, we introduce BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video. BuddyVQA contains 21.6K questions linked to 6K highlight moments across 1,012 long, egocentric videos. It features two key characteristics that are common in daily first-person QA assistance but are largely overlooked in existing VideoQA benchmarks: ego-deictic expressions and interactively chained questions (e.g.,"Where is it?","How to get there?"). These require models to infer a user's in-situation intent by resolving visual pronouns in the context of egocentric visual and QA contents, with both grounded in a long-form streaming setting. To tackle the challenges, we propose MyBuddy, a companion-style QA assistant that highlights a multimodal chain-of-thought reasoning mechanism to infer the final answer based on the historical QA and visual content. An additional question filter and multi-level memory are designed to facilitate efficient QA and visual information retrieval under streaming QA settings. Experiments show that MyBuddy significantly enhances the performance of foundation models on BuddyVQA. Moreover, these gains generalize to other streaming and common video QA benchmarks, demonstrating the applicability and effectiveness of our approach. Our code and dataset are available at https://github.com/QHUni/BuddyVQA
The Identity-conditioned Queries task is introduced, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges.
Shibo Gao, Chongxiao Wang, Chenglong Huang et al.· 0 citations
This work proposes a novel task named Egocentric Video Dialogue, which refers to answering questions by considering both the video content and dialogue history, and designs a new dataset, EgoDialogue, which constructs continuous dialogues based on video question answering datasets.
Jia-Yi Zou, Chaofan CHEN· ACM Transactions on Multimed...· 0 citations
This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill...
Online information is increasingly consumed in video format. Much of this comes in the form of *raw video*: continuous footage taken on a cell phone, with a hand-held camera, or via CCTV, which is then directly uploaded to social media platforms and content sharing services. Whereas professional or even amateur-edited...
Reno Kriz, David Etter, Alexander Martin et al.· 0 citations
The system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass, and reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters.
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning...
Shu-Lin Tian, Junsu Kim, Shuai Liu et al.· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.