Fine-tuning LLaVA-1.6-7B on taxonomy-derived labels for ACT and CLARIFY indicates the feasibility of training vision-language models using annotations from the authors' taxonomy, and motivates a constrained protocol for more consistent dialogue collection.
Abstract
Robots that follow natural-language instructions in everyday indoor environments must act on incomplete human utterances. Instructions often omit essential information, such as the identity of an out-of-view object, an intended destination, or the user's goal. Existing datasets contain little real-world situated dialogue and provide few practice-grounded criteria for deciding when a robot should act, confirm, clarify, or refuse. We retrospectively analyze a pilot Wizard-of-Oz study in which five participants performed everyday indoor tasks, including door opening, drawer opening, feeding, drinking, and cleaning, with a wheelchair-mounted mobile manipulator while the wizard responded without a formal communication policy. This preserved authentic user behavior but produced inconsistent robot-side decisions, motivating an explicit decision scheme. From 40 episodes, we derived a hierarchical taxonomy of six response modes (ANSWER, REPORT_DONE, REFUSE, CONFIRM, CLARIFY, ACT) and four ambiguity types (intent, referential, spatial, intelligibility). Two human annotators and an AI annotator applied the scheme to the pilot data. Clean-label rates were 91% and 89%, and Cohen's ranged from 0.72 to 0.95 across decision-point, mode, and ambiguity levels for both human-human and human-AI comparisons. Fine-tuning LLaVA-1.6-7B on taxonomy-derived labels for ACT and CLARIFY indicates the feasibility of training vision-language models using annotations from our taxonomy. Remaining boundary cases in decision-point identification and REPORT_DONE motivate a constrained protocol for more consistent dialogue collection.
A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt a person. Choosing well requires knowing how much the robot's sensors reveal about the cause and how reliable the robot's own diagnosis is. We build a simulated benchmark...
SAIN is presented, a zero-shot framework that turns active dialogue into persistent navigation state and supports dialogue-to-state conversion as an effective zero-shot mechanism for long-horizon interactive instance navigation.
Robots operating in human-centered environments are typically designed to execute explicit instructions, and most robot-learning datasets likewise pair observations with task instructions or low-level actions. Although recent work has begun to explore proactive embodied assistance, existing resources target different s...
Zhi-Hao Gu, Kenny Zhu, Yuan-Feng Wu et al.· 0 citations
Adaptive behavior is a fundamental requirement for social robots operating in real-world environments, where interactions must dynamically respond to both observable actions and inferred user states. In this work, we propose a novel framework that models human–robot interaction as a planning problem, where adaptive con...
Giulia Berettieri, Anna Allegra Bixio, Lucrezia Grassi et al.· Companion Publication of the...· 0 citations
Although gesture-based explanations for robotic behavior seem promising for enhancing the transparency of robotic systems, careful design is crucial to their success. Based on a pilot laboratory study (
N
= 42) that highlights the need for a systematic framework for gesture-based design of transparent robotic syste...
I. Hein, Daniel Ullrich, Christina Brunner et al.· Frontiers of Computer Scienc...· 0 citations
Successful conversations require speakers to align on conceptual understanding, a challenging but crucial task in human-robot interaction. With the increasing use of large language models for dialogue, robots move from passively acquiring human conceptualizations to actively shaping alignment. However, the design space...
Sheng-Chen Zhang, Mei-Ying Li, Zi-Xuan Wang et al.· Proceedings of the 14th Nord...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.