Autonomous mobile robots require reliable coordination among navigation, perception, tracking, and precision approach modules to complete indoor object-search missions. Existing systems often remain fragmented, treating navigation, detection, tracking, and docking as separate tasks rather than as an end-to-end pipeline. This study has two objectives: to validate a fault-tolerant coordination architecture for autonomous search-and-approach behavior and to compare two search strategies under controlled indoor target-position scenarios. The primary contribution is a methodological integration framework based on a Finite State Machine (FSM) that coordinates ROS 2 Navigation2 global navigation, YOLO11n object detection, centroid tracking, and Image-Based Visual Servoing (IBVS), while managing transitions among navigation, visual servoing, recovery, and mission-completion states. A quantitative Gazebo simulation experiment used 40 controlled trials to compare Random Exploration and Waypoint-Based Search. The integrated system achieved a 100% mission success rate without command conflicts, indicating effective FSM-based coordination between global navigation and local visual control. Waypoint-Based Search was more efficient when the target was aligned with predefined nodes, achieving a mean detection time of 50.55 s compared with 164.80 s for Random Exploration. Conversely, Random Exploration performed better when the target was away from predefined paths, reducing mean detection time to 87.00 s compared with 183.64 s. Fault-tolerant behavior was demonstrated in simulation through successful mission completion despite repeated LiDAR-triggered obstacle-recovery events during visual approach. These findings show that search efficiency depends on alignment between exploration design and spatial structure, not universal strategy superiority.
Human Activity Recognition (HAR) from smartphone accelerometer data is widely studied on the WISDM dataset, but random sample-based partitions can allow windows from the same subject to appear in both training and evaluation data, producing optimistic estimates of cross-subject generalization. This paper investigates the DeepConvLSTM architecture on WISDM v1.1 using only tri-axial smartphone accelerometer signals under a subject-disjoint split comprising 25 training subjects and an 11-subject evaluation partition. Nine controlled experiments varied window configuration, model capacity, temporal pooling, and augmentation strategy. Under this single fixed subject split, without repeated random seeds or statistical comparisons, the best-performing configuration among the nine experiments used a 60-sample window (3 s) with 50% overlap and on-the-fly jitter and scaling augmentation, achieving 90.12% accuracy, compared with 89.12% for offline augmentation, 88.75% for class-specific augmentation, and 87.49% with label smoothing. Global Average Pooling showed an observed class-level trade-off relative to last-timestep pooling in the evaluated comparison rather than a general architectural advantage. A persistent 8–10 percentage-point training–evaluation gap remained, with notable confusion among stair-related locomotion classes, which may partly reflect limited subject diversity, class imbalance, and accelerometer-only sensing. Importantly, the same 11-subject evaluation partition was consulted for early stopping, learning-rate scheduling, comparison of all nine experiments, and final model selection; therefore, the reported 90.12% accuracy should not be interpreted as performance on a fully untouched test set. These findings provide a leakage-resistant but selection-sensitive benchmark for subject-independent DeepConvLSTM evaluation on WISDM.
Fakhrul Zidan Nurrohman, C. Dewa· bit-Tech· 0 citations