Skip to content
Preprint

CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

CAP is proposed, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information.

Abstract

Humanoid locomotion across complex terrain demands forward-looking exteroception to anticipate obstacles, yet this signal is unreliable in real-world deployment, failing partially and intermittently. Existing perceptive policies often assume that depth observations remain clean and in-distribution, while recent attempts to unify perceptive and blind control typically route or switch between separate sub-policies, leaving recoverable information in partially corrupted depth unexploited. We instead propose CAP, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information. A coupled training recipe pairs a depth-noise curriculum on the world-model input with world-model feature dropout on the policy-facing latent, exposing the policy to failures across the entire perception-quality spectrum. In simulation, CAP matches or improves upon perceptive baselines when depth remains informative, and degrades more smoothly than a binary-switching baseline as perception worsens. On the Unitree G1, controlled trials and indoor-outdoor deployments demonstrate perception-robust locomotion under intermittent occlusion, real-sensor corruption, and outdoor depth artifacts.

View source

Similar papers

Preprint Sep 2026

REDACT: Robust Perceptive Locomotion under Unseen Visual Corruption

Depth-conditioned locomotion policies have demonstrated impressive agile maneuvers, but can be steered to unpredictable actions when observations are outside their training distribution. Occlusion, invalid returns, sensor noise, and visual distractors can shift deployment observations away from nominal simulated depth....

Natapat Kirdwichai, Tobias Driskell-Poole, Andrei Sontea et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models

Vision-based legged locomotion methods assume clean depth at training time and rely on hand-tuned post-processing filters at deployment. However, filter parameters are rarely disclosed, hindering reproducibility, and performance degrades substantially when depth noise is left unaddressed. Building noise robustness dire...

Yo-Han Choi, Min-Jun Kim, Jin-Sung Kim et al. · 1 citation
Preprint Aug 2026

ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation

ViTaR is introduced, which reframes tactile feedback from an action-generating perceptual input to an execution modulator that selects and scales bounded residual corrections atop a frozen VLA, preserving pretrained capabilities by construction.

Yi Wang, Renjun Wu, Jin-Yan Liu et al. · 0 citations
Preprint Sep 2026

ViBe: Visual Behavior Adaptation for Perceptive Humanoid Whole-Body Control

Motion tracking provides a scalable recipe for humanoid whole-body control. By design, the resulting trackers lack exteroceptive feedback hence reacting to the environment remains the responsibility of a higher-level planner. Existing perceptive controllers train geometry-only encoders from scratch, trading semantics f...

Lokesh Krishna, S. Venkatesan, An Zhang et al. · 0 citations
Preprint Sep 2026

ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning

Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each nove...

Chiyoung Kim, Min-Sung Choi, Jin-Ho Ju et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.