A novel egocentric and exocentric video dataset capturing real-world collaboration in cooking scenarios, and establishes benchmarks for Joint Attention Estimation, Socially Conditioned Object Interaction Anticipation, and Collaborative Handover Prediction, enabling research on multimodal perception, proactive assistance, and collaborative planning.
Abstract
Human-human collaboration is a fundamental aspect of everyday life, essential to success in a wide range of goal-directed activities from household tasks to professional teamwork. While much research has focused on modeling coordination and task execution, the cognitive processes that support such collaboration, particularly Theory of Mind (the ability to infer the mental states of others), remain difficult to study in natural settings. To address this gap, we introduce a novel egocentric and exocentric video dataset capturing real-world collaboration in cooking scenarios. The dataset integrates multi-perspective video, high-quality audio, gaze tracking, and 3D scene and object scans, with annotations for shared attention to objects, social cues and interactions between agents, as well as agent-object interactions. We establish benchmarks for Joint Attention Estimation, Socially Conditioned Object Interaction Anticipation, and Collaborative Handover Prediction, enabling research on multimodal perception, proactive assistance, and collaborative planning. By providing temporally aligned, richly annotated multimodal data, CoMind facilitates the development and evaluation of AI systems capable of modeling complex social interactions and reasoning about human behaviors in collaborative environments. Our dataset and benchmarks are made available at https://comind.ethz.ch/.
This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill learning, human-to-robot transfer, and embodied decision making.
Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human-centered visual intelligence is especially demanding because it studies people as expressive, socially situated subjects whose meaning is rarely conveyed by appearance alone. It couples vision with audio and language across four representative tasks: human emotion recognition, human video generation, human voice cloning, and human video matting. Yet existing resources remain task-specific, providing modalities and annotations for individual problems rather than a shared foundation coordinating understanding and generation. This limits multimodal signal use and broader research. We address this gap with HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence. It contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alpha mattes. We evaluate diverse open- and closed-source models across the four tasks under a unified zero-shot protocol using automatic metrics, criterion-specific mean opinion scores, and multiple cross-task analyses. Results show that (i) linguistic content dominates current emotion recognition, while purely visual affect recognition is weakest; (ii) in video generation and voice cloning, automatic metrics and human judgment agree overall but differ in their top rankings, requiring joint reporting; (iii) boundary fidelity under motion is the main remaining obstacle for human matting; and (iv) task difficulty varies across emotions, models, and metrics, with notable cross-task correlations. The dataset and results are available at https://github.com/GML-MMGroup/HUG-VIS.
Fei Ma, Zebang Cheng, Ming-Hui Li et al.· 0 citations
Collecting contact-rich robot experiences at scale remains a major bottleneck for generalizable manipulation. Beyond data quantity, robot learning also requires diverse experiences across embodiments, viewpoints, and scenes. Human egocentric videos provide abundant physical interactions, but each video captures only a narrow slice of experience under a single body, camera trajectory, and environment. We propose AnyWorld, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations. Our model factorizes an interaction into action, camera, and embodiment: action controls capture the motion structure, camera controls specify viewpoint evolution, and the target embodiment context defines the acting body and its interaction geometry. This formulation enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain experiences while preserving the underlying dynamics and object interactions. We train the model with large-scale human interaction pretraining followed by mixed-embodiment fine-tuning. Experiments show that our model supports controllable recomposition across embodiments, viewpoints, and scenes, and we further demonstrate that the generated data can improve manipulation performance on the RoboCasa GR1 tabletop benchmark and a real IRON humanoid robot. Beyond aggregate gains, we test whether unpaired human experience can be recomposed into robot-native video-action pairs that target a policy gap. Controlled IRON interventions correct a spurious completion prior and establish language-grounded spatial target selection; an action-only counterfactual intervention fails to learn the latter reliably, showing that both action calibration and visual recomposition are necessary.
Cheng Chen, J. Bai, Jiacheng Wei et al.· 0 citations
Conversational User Interfaces (CUIs) are rapidly advancing, moving from single-task assistants to powerful and engaging artificial agents. As CUIs become embedded in daily life, the implications for users interacting with such systems demand closer scrutiny. While some issues are immediately identifiable (e.g. privacy risks, misinformation, AI hallucinations), others may only become visible after longer periods of use, including over-reliance, erosion of human agency, or the normalization of biased or exclusionary language. This workshop aims to gather a multidisciplinary community to reflect on pressing challenges and uncertainties and explore risk analysis and mitigation strategies in CUI design and deployment. Through presentations, discussions, and collaborative design activities, participants will examine immediate and longitudinal risks and share empirical insights. The activities will support developing interdisciplinary frameworks to understand and handle unintended consequences of CUIs. The workshop seeks to build an international network of scholars and practitioners to promote responsible and human-centred conversational AI.
Manveer Kalirai, C. Wei, Thomas Essmeyer et al.· International Conference on...· 0 citations
A central challenge in understanding joint action is explaining how individuals achieve successful coordination in dynamic, real-world settings. Although coordination is thought to depend on anticipating future events, including the actions of others, the precise contribution of such predictive processing remains unclear, since most existing evidence comes from simplified laboratory tasks that infer anticipation indirectly from reaction times. To address this gap, we studied 70 participants (forming 64 pairs) during solo and joint sessions of a turn-taking ball-hitting task while recording multimodal behavioral, physiological, demographic, and social measures. We found that successful coordination was most strongly associated with anticipation of one’s own action outcomes, outperforming physiological synchrony, motor behavior, demographic characteristics, and social closeness. Partners whose predictions of their own action outcomes were more closely aligned coordinated better, likely because each behaved in ways the other implicitly expected. A Bayesian generative model further showed that well-coordinated pairs relied more heavily on prior expectations than on incoming sensory evidence when generating predictions during joint action. These results suggest that efficient coordination in dynamic real-world tasks emerges primarily when individuals’ internal models for predicting their own actions are well aligned across partners. Our quantitative multimodal approach provides a framework for disentangling the contributions of predictive, physiological, motor, and social factors to coordination in naturalistic interactions. Combining eye-movement, body-kinematics, heart-rate, and individual-characteristic data with machine learning, this study shows that coordination during dynamic motor interaction reflects how similarly partners predict their own actions.
P. Putra, Fumihiro Kano· Communications Psychology· 0 citations
Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-world collaboration fundamentally requires coordination under shared agency, including intent understanding, temporal synchronization, protocol adherence, and safe interaction in dynamic environments. To address this gap, we introduce HRIBench, a diagnostic benchmark for intent-aware human-robot collaboration based on executable interaction scenarios. HRIBench represents collaborative tasks as structured scenario scripts that explicitly model agent roles, temporal dependencies, coordination constraints, and human behavior distributions. Building on this abstraction, HRIBench defines three representative interaction roles: Instructor, Collaborator, and Intruder, covering intent communication, joint coordination, and robustness under human intervention. The benchmark contains 13 role-conditioned tasks with over 650 evaluation episodes generated from diverse interaction trajectories and scene variations. Beyond binary task success, HRIBench introduces interpretable interaction-centric metrics spanning synchronization, responsiveness, protocol compliance, and safety. We evaluate adapted policies based on GR00T, pi0.5, and ACT under a unified protocol. Results show that current foundation robot policies struggle substantially in collaborative settings despite strong manipulation ability, revealing major limitations in temporal coordination and intent-aware behavior. Fine-tuning on HRIBench consistently improves collaborative performance. In a real-world adaptation study, simulation data generated by HRIBench improves GR00T N1.5's physical-task success rate from 0.10 to 0.43, demonstrating the benchmark's value for advancing interaction-centric robot learning.
Chang Liu, Jiawei Zhang, Tao Zhang et al.· 0 citations