HuMiT is presented, a whole-body teleoperation system built on a minimal reference target that requires only a minimal reference target, consisting only of root height, root velocity, and sparse keypoint positions at the current frame, yet achieves competitive or superior tracking superiority compared to methods relying on more diverse reference states.
We present a real-time upper-body human-to-humanoid motion imitation framework driven by neuromorphic event-based vision. This work addresses practical perceptual bottlenecks of conventional frame-based RGB sensors, specifically their difficulty in high dynamic range (HDR) scenes and rapid motions due to fixed integration times. By leveraging the Prophesee EVK4 event camera, which operates asynchronously with high temporal resolution and a dynamic range exceeding 120 dB, our system supports stable tracking in conditions where standard vision pipelines degrade, such as severe backlighting and very low light environments below 5 lux. The architecture integrates a low-latency Perception Module, utilizing optimized event accumulation and gravity-aligned inertial fusion, with a causal Motion Module (TWIST) that performs online kinematic retargeting. We validate the system on an embedded NVIDIA Booster T1 platform and an 18-DoF humanoid upper-body setup, demonstrating an end-to-end photon-to-action latency of 23-34 ms and advantages over RGB baselines under our experimental setup. The results indicate a practical trade-off: events can be preferable for fast or poorly lit upper-body teleoperation, whereas well-lit static scenes may favor RGB or hybrid sensing.
Humanoid teleoperation for demonstration collection requires coordinated whole-body motion, continuous dexterous hand control, and viewpoint control. Existing systems either simplify hand commands or depend on dedicated wearable sensors for fine-grained hand motion. We introduce Teleopit, a full-embodiment teleoperation system that maps body, hand, and head signals from VR to a humanoid body, configurable dexterous hands, and a 2-DoF active vision module. A history encoder and failure-aware rewind sampling improve the motion tracker on both motion-capture and live VR references. An optimization-based hand retargeter combines normalized finger directions, fingertip closure, and thumb-frame alignment to map human hand motion to different dexterous hands without tuning hand-specific objective or solver hyperparameters. Component experiments evaluate tracking success rate and retargeting behavior, while real-robot teleoperation demonstrates coordinated locomotion, manipulation, and viewpoint control. ACT and GR00T N1.7 policies trained on 96 successful demonstrations collected with Teleopit achieve task success rates of 90.0% and 95.0%, respectively, when deployed on the humanoid. The project page is available at https://botrunner64.github.io/teleopit-page.
Bingqian Wu, Zichen Xu, Xianghui Fan et al.· 0 citations
Humans routinely wield tools, swap grasps, and reposition objects within a single hand, seamlessly orchestrating contact transitions that span translation, reorientation, and finger gaiting. Endowing robot dexterous hands with this level of in-hand dexterity through teleoperation requires precise control of object motion via dynamic hand-object contact, yet current teleoperation systems remain far from this capability. To bridge this gap, we take a major step towards human-level dexterous teleoperation by introducing TeleDexter, a hand-object co-tracking controller that maps operator intent into learned, low-level contact execution. The controller is trained on consecutive co-tracking subgoals derived from human reference motions, utilizing a hybrid reward that couples sparse subgoal objectives with dense tracking rewards to enable learning across diverse interaction modalities rather than frame-wise trajectory imitation. The entire pipeline requires only single-stage RL and, with random action masking and domain randomization, transfers zero-shot to the real robot. We evaluate TeleDexter on seven challenging dexterous teleoperation tasks spanning object reorientation and long-horizon tool use across two dexterous hands, achieving a 75% average success rate where all baselines consistently fail. Furthermore, the collected demonstrations successfully train autonomous policies via behavioral cloning, marking a concrete step towards human-level dexterous teleoperation.
Puhao Li, Zeyuan Chen, Yingying Wu et al.· 0 citations
This work introduces HumanTracker, a preference-aligned metric trained on 12K motion pairs containing 24K motions that better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Dai-En Liu, Zekun Qi, Jiayu Zeng et al.· 0 citations
Real-world humanoid tasks involve physical interaction with objects and humans, yet current controllers either reject external forces as disturbances or restrict compliance to limited body links while ignoring angular effects. We present LAC, a general whole-body controller that simultaneously realizes commanded Linear and Angular Compliance for wrenches applied to the upper body. First, we synthesize whole-body compliant responses into a large-scale augmented dataset. Sampled force and couple events are imposed on contact frames extracted from human interaction data. At each contact link, the external force and a virtual torque from the passively yielding kinematic chain drive a virtual admittance under the commanded stiffness. Subsequently, teacher-student reinforcement learning trains a single policy to track the compliant motions under external wrenches. Finally, extensive simulation and real-world experiments demonstrate whole-body compliant responses to wrenches across the upper body, monotonic modulation over the full range of both stiffness commands, and applicability to teleoperated loco-manipulation tasks. Project website: https://lac-humanoid.github.io/
Yang Liu, Zhongkai Gu, Wei Zhu et al.· 0 citations
Tendon-cable transmission can reduce distal-limb inertia in full-size humanoids, but its elasticity, hysteresis, backlash, and multi-joint coupling introduce state-dependent joint-to-motor discrepancies. We present a hierarchical whole-body tracking framework for the 28-DoF Droid X3 that separates high-level motion learning from transmission compensation. A reference-residual policy is trained in simulation by single-stage proximal policy optimization (PPO) using a unified robot-space motion representation, globally anchored tracking rewards, hierarchical hard-example sampling, and tendon-oriented domain randomization. In simulation checkpoint evaluation, more than 90% of 12,674 tested reference motions are completed. Independently, a state-conditioned mapper is trained offline through a differentiable motor–joint forward model identified from physical motor-excitation data and connected in series between the frozen policy and the low-level motor controller. Randomized repeated Mapping-OFF/ON trials are conducted on two nominally identical Droid X3 units. Within every robot–motion block, the frozen PPO checkpoint, reference trajectory, controller settings, safety bounds, and frozen mapper weights are held fixed; complete trials are the statistical units. OFF converts desired joint positions with the robot-specific fixed static calibration, whereas ON feeds the complete policy-level desired-joint vector and measured plant state to the frozen mapper, which directly outputs the complete motor-position command. Across the complete physical trials, the aggregate action-completion rate is 68% with Mapping OFF and 79% with Mapping ON, an increase of 11 percentage points. Representative walk, squat, and dance trajectories illustrate lower tracking errors under Mapping ON, while individual frames and selected temporal fragments are used only for visualization.
Wencong Gan, Jie-Hui Chen, Qingdu Li et al.· Biomimetics· 0 citations