Jun 2026
LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music
A novel framework that leverages the reasoning capabilities of Large Language Models to synthesize complex robotic actions from a rich tapestry of multimodal human inputs: natural speech, hand gestures, and music/sound beats.
Snehasis Banerjee, R. Dasgupta
· arXiv.org · 0 citations