Diffusion-Driven Point Cloud Adaption for Depth-Only 6D Pose Estimation in Cluttered Robotics Bin-Picking
D pose estimation in densely-packed, industrial environments remains a challenging perception problem due to severe occlusions, sensor noise, and the fragmentary nature of reconstructed 3D observations. To this end, we present a generative module for holistic, depth-only 6D pose estimation in cluttered bin-picking that integrates a conditioned Denoising Diffusion Probabilistic Model (DDPM) as a learned data-centric inference adaptor for point cloud segments. Building on a holistic staged-heatmap fusion backbone that produces focused, memoryefficient point cloud fragments around candidate objects, we show that sharpening and completing these fragments with a diffusion-based point cloud autoencoder substantially improves downstream per-point voting pose estimation. Our method integrates a velocity-targeted DDPM, implemented via a Point Transformer V3 (PTv3) backbone, to reconstruct sharper and more complete object-level geometry from imperfect scene fragments. The denoiser is trained to explicitly align geometric completion with the 6D pose objective. At inference we apply a lightweight sampling schedule to produce a single sharpened segment per object candidate, preserving time-efficiency for robotics bin-picking. With pilot experiments on the IPD dataset, we demonstrate that diffusion-based completion reduces pose error and increases robustness to occlusion, sensor noise, and severe fragmentation, while retaining the holistic advantages of scene-level reasoning. We argue that diffusion models are a promising direction for scene denoising and completion in 6D pose estimation and provide a practical integration strategy for robotic perception systems.