Experiments show that M-VTOP achieves sub-millimeter accuracy under complex geometries, occlusions, and tight tolerances, demonstrating its promise for high-precision robotic manipulation.
A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.
Hui Zhang, Yue Wang, Kang An et al.· Signal, Image and Video Proc...· 0 citations
This study investigates in-hand pose estimation of a USB stick that is already held within a robotic gripper and provides a controlled task-specific analysis showing that structured model selection can improve multimodal grasp-pose regression within the evaluated setup.
A. Altenbuchner, Bsher Karbouj, Fabian Dilly et al.· IEEE Access· 0 citations
A unified monocular vision-based grasping framework that targets both soft and rigid objects within a single control pipeline, using only RGB input and a position-controlled gripper, and is validated in real-world pick-and-place experiments.
Shail V Jadav, Dongheui Lee· 2026 IEEE/ASME International...· 0 citations
This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.
Xinmiao Du· Poster Volume 0007 The 2026...· 0 citations
For robotic dynamic grasping of moving objects in conveyor-belt scenarios, accurate and robust 6D pose estimation and tracking are essential for reliable grasping. However, existing deep-learning-based methods usually rely on large amounts of supervised data for specific objects or categories, which limits their generalization, deployment efficiency, and flexibility for rapid object changeover in industrial applications. To address these challenges, this paper proposes FreeTrack6D, a training-free unified segmentation and 6D pose tracking framework. Relying only on the CAD model of the target object, FreeTrack6D can be applied to dynamic tracking and grasping of unseen objects without additional object-specific training. Specifically, an adaptive multi-cue mask generation module is first introduced to generate frame-wise target masks in real time, which provides target-region constraints for initial pose registration and subsequent pose refinement. This helps reduce the influence of background interference and target-region misalignment caused by rapid motion. Based on the generated mask, RGB-D observations, and the CAD model, FoundationPose is used for initial 6D pose registration and subsequent pose refinement. To improve tracking robustness under large inter-frame motion and rotational variations, a Kalman-guided multi-hypothesis refinement strategy is further designed, where multiple candidate poses predicted from historical motion states are refined and selected according to mask consistency. In addition, a Pose Consistency-aware Association and Gating mechanism is developed to reject abnormal detections, protect the filter state, and trigger re-initialization when consecutive mismatches occur. By integrating frame-wise mask generation, multi-hypothesis pose refinement, motion prediction, observation gating, and visual-servo-based robot control, FreeTrack6D forms a closed-loop training-free dynamic grasping framework.
Zongwang Han, Long Chen, Shiqi Wu· Engineering Research Express· 0 citations
Accurate object localization is essential for enabling autonomous operation of indoor robotic systems. As a low-cost, compact, and flexibly deployable solution, monocular visual object localization (MVOL) is highly applicable to lightweight embedded robotic platforms. However, conventional data-driven MVOL methods suffer from inherent limitations of 2D-to-3D ill-posed mapping, which is caused by the incapability of constraining the spatial physical logic of real scenes, resulting in severe depth ambiguity, inaccurate scale estimation, and physically unreasonable predictions. To address these issues, this paper proposes a novel physics-guided monocular visual localization framework termed PC-IMVL for indoor scenarios. The PC-IMVL integrates deep visual perception with embedded physical modeling, which explicitly introduces spatial physical constraints into the network optimization process and builds a physical consistency-aware loss function to regularize 3D position and pose estimation. Combined with a lightweight tailored architecture, the framework enables efficient and reliable embedded deployment. Offline experiments and real-world online tests validate the effectiveness of the proposed method. PC-IMVL yields average absolute errors (AE) of 0.095–0.333 m, reducing the localization error of early fusion methods by more than 50%. Within a working distance of 3–4 m, it achieves a relative error (RE) of 2.4% and a horizontal viewing angle error (VAE) below 2°, outperforming existing state-of-the-art MVOL methods. The effectiveness of the physical guidance mechanism is verified. This work provides a practical high-precision localization solution for embedded indoor robotic systems.
Haorui Ge, Luzheng Bi, Weijie Fei et al.· Italian National Conference...· 0 citations