Skip to content
Conference

Pixel-Level Grasp Pose Measurement in Cluttered Scenes via Sim2Real Domain Randomization

Jul 2026 · 2026 IEEE 27th China Conference on System Simulation Technology and its Applications (CCSSTA) · pp. 139-144 · 0 citations · 22 references
View source

Similar papers

#artificial intelligence Preprint Aug 2026

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

A reinforcement learning-based framework for robotic grasp refinement, integrating keypoint-based object representations with a Deep Q-Network (DQN), is proposed, offering a scalable and adaptable solution for contact-rich manipulation tasks.

Amir Arsalan Nematollahi, Shayan Ahmadi, M. T. Masouleh et al. · 0 citations
Open access 2026

Grasp Pose Estimation of Articulated Objects Based on Semantic and Geometric Feature Fusion

A deep learning-based grasp estimation model designed to enable robotic manipulation with articulated objects that incorporates the attention-based semantic and geometric feature fusion (ASGF) module improved the grasp success rate in the evaluated setting.

Dongwoo Lee, Yeongmin Kim, Seong-Bo Jo et al. · 0 citations
2026

GFLA: A Grasping Framework With Learning-Based Perception and Analytical Modeling for Single-View Scenes

Antipodal grasping from single-view red-green-blue and depth (RGB-D) images is challenged by occlusion and partial observability, making purely analytical inference ill-posed. We present the Grasping Framework with Learning-Based Perception and Analytical Modeling (GFLA), which fuses learning-based perception with analytical modeling. GFLA projects antipodal contacts to the image plane, samples grasp candidates via inverse projection, and ranks them with a force-closure metric. To compensate for the information loss inherent in single-view observations, we introduce two grasping hypothesis-guided modules: 1) a contact projection detection network that localizes graspable regions and predicts antipodal projections on visible surfaces, and 2) a 3-D U-Net-based scene completion network that completes geometry and provides explicit collision cues. On GraspNet-1Billion, GFLA achieves its largest improvement on the novel object set (average precision (AP) 35.88%, an improvement of 7.59%), demonstrating superior generalization to previously unseen object categories while also attaining a competitive overall AP of 57.84% (an improvement of 1.33%). Real-robot experiments in cluttered environments, without domain adaptation or fine-tuning, achieve grasp success rates of 95.42% for single-object scenes and 90.12% for multiobject scenes, demonstrating strong practical robustness.

Xiao Ning, Jianzhong Yang, Si Huang et al. · 0 citations
Aug 2026

Model-agnostic pose estimation for enhanced collaborative robot grasping via binocular vision

A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.

Hui Zhang, Yue Wang, Kang An et al. · 0 citations
Open access Jul 2026

WaveletkAN Pose Refinement for Robot Bin Picking Driven by Cluttered RGB-D Point Clouds

Grasping cluttered boxes is still limited by the gap between rough 6-degree-of-freedom perception and the millimeter-level alignment required for robot execution. WaveletkAN is a pose-refinement model based on RGB-D point clouds proposed in this paper to correct object pose hypotheses for dense bins with occlusion, specular depth noise, and self-similar industrial parts. The method represents each candidate as a local observed point set, a CAD-derived canonical set, and a compact bin-context tensor, then predicts a residual rigid motion via Kolmogorov-Arnold network layers parametrized by learnable wavelet atoms. A confidence-weighted correspondence field and a residual SE (3) update can be used with the traditional proposal generator and grasp planner. WaveletkAN reduced the median translation error from 7.8 mm to 2.9 mm and the median rotation error from 5.6 degrees to 1.9 degrees after refinement in a simulated-real mixed benchmark with 18 object categories and 42,600 evaluated hypotheses. The successful pick rate in dense clutter rose to 93.1 %, and the mean refinement latency of the industrial GPU was still 11.6 ms. Ablation experiments showed that removing the wavelet basis increased ADD-S by 31.7%, and omitting context gating reduced top-1 executable pose recall by 6.8 %. Based on the above experiments, multi-scale functional parameterization can achieve robust and deployable pose correction for RGB-D robotic bin-picking systems.

Charalampos Evangelou, Iakovos Maniatis · 0 citations