Skip to content
Conference

Voice-Controlled Robotic Arm System for Tabletop Manipulation via Large Language Model and 3D Vision

Jul 2026 · 2026 23rd International Conference on Ubiquitous Robots (UR) · pp. 15-20 · 0 citations · 12 references

Abstract

Natural language interfaces can lower the expertise barrier for operating robotic manipulators by allowing users to express goals in everyday speech. This paper presents RANLP-Arm, a modular bilingual voice-to-manipulation system that converts spoken Thai or English instructions into real-time robotic arm actions for tabletop pick-and-place tasks. The system integrates Gemini-based live voice interaction with function calling, stereo 3D object perception and segmentation, persistent object tracking with EMA smoothing and coordinate locking, camera-to-robot calibration with optional IDW residual correction, and TCP/IP robot control in a unified Python architecture for real-time operation. We evaluate the system on a physical robot using 100 pick-and-place trials and 40 bilingual voice-commanded trials. The affine calibration model achieves a mean correspondence error of 3.72 mm on 13 calibration points, while end-to-end experiments achieve 82.0% task completion, 100.0% intent recognition, and 95.0% voice-to-task success. All observed failures were caused by grasp instability rather than language understanding, perception, or calibration errors. These results show that bilingual voice-driven manipulation is practical for tabletop pick-and-place, while the main remaining limitation lies in end-effector robustness.

View source