Edge-deployable greenhouse tomato cluster harvesting robot integrating YOLOv8n-BiFPN-WIoU and ROS-based autonomous control
Abstract
To develop an edge-deployable greenhouse tomato cluster harvesting robot capable of accurate fruit detection, three-dimensional localization, autonomous navigation, and robotic harvesting under complex greenhouse conditions. The objective was to improve detection performance under fruit occlusion and variable illumination while maintaining real-time inference and stable robotic operation on embedded computing hardware. A lightweight YOLOv8n-based tomato detection model, termed YOLOv8n-BiFPN-WIoU, was developed by incorporating a bidirectional feature pyramid network and Wise-IoU loss to improve multiscale feature fusion and localization accuracy. An RGB-D vision system based on the Intel RealSense D435 was used for three-dimensional fruit localization. The robotic system integrated a rail-guided mobile platform, a six-axis manipulator, a dual-claw harvesting end-effector, and a ROS-based distributed control architecture. The vision model was optimized and deployed on an NVIDIA Jetson Nano using ONNX and TensorRT. Robot kinematics, hand-eye transformation, trajectory planning, and autonomous harvesting procedures were integrated to enable station-to-station operation and fruit harvesting. The proposed YOLOv8n-BiFPN-WIoU model achieved a precision of 88.534%, recall of 89.377%, F1-score of 88.954%, mAP@0.5 of 92.338%, and mAP@0.5:0.95 of 71.368%. Compared with the baseline YOLOv8n model, these metrics improved by 2.143, 2.779, 2.460, 4.587, and 5.675 percentage points, respectively. The proposed model increased the parameter count from 3.2 million to 3.6 million and computational complexity from 19.6 to 22.4 GFLOPs. After ONNX and TensorRT optimization, the model achieved an inference speed of 21.1 frames per second at an input resolution of 640 × 640 with batch size 1 on the Jetson Nano. The robotic system achieved a mean positioning error of 3.70 ± 0.32 mm across five representative targets, demonstrating the feasibility of accurate visual localization and robotic positioning for greenhouse tomato harvesting. The proposed edge-deployable greenhouse tomato cluster harvesting robot integrates lightweight deep learning, RGB-D three-dimensional localization, ROS-based autonomous control, rail-guided mobile operation, and robotic harvesting into a unified system. The YOLOv8n-BiFPN-WIoU model improved tomato detection performance while maintaining real-time inference on embedded hardware, and the integrated vision and manipulation system demonstrated accurate positioning suitable for autonomous harvesting. The results indicate that the proposed architecture provides a practical foundation for intelligent greenhouse tomato harvesting under complex visual and operational conditions.