Jul 2026· 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM)· pp. 1-6· 0 citations· 14 references
Abstract
This paper proposes a domain adaptation method for different lidars in point cloud semantic segmentation using semi-supervised learning and map annotation. This approach addresses the significant costs associated with acquiring point-level annotations in target domains. We first perform representation learning using large-scale unlabeled data. Subsequently, we conduct classification learning using a point cloud map integrated via coordinate transformations, eliminating the need to annotate individual frames. Experiments on our original TC dataset demonstrate that our method effectively bridges the domain gap from the SemanticKITTI dataset, improving the mean Intersection over Union (mIoU) by 48.6 points over a zero-shot baseline. Furthermore, our approach outperforms a strong baseline trained from scratch on the target domain by 7.2 points in mIoU. These results establish a cost-efficient solution for deploying diverse sensors.
A lightweight boundary-aware learning framework that explicitly models boundary regions during training is proposed, showing that incorporating boundary-aware supervision provides an effective and efficient approach to improving segmentation quality in challenging regions.
Waseem Iqbal, J. Paffenholz· The International Archives o...· 0 citations
A SAM-guided framework for point cloud oversegmentation that significantly improves boundary recall and maintains high oracle accuracy while maintaining high oracle accuracy, and generalizes well to unseen datasets without retraining, showing strong cross-dataset inference capability.
Dening Lu, Michael A. Chapman, Jonathan Li· The International Archives o...· 0 citations
Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense annotation. This motivates a simple question of whether high-quality LiDAR training data can be produced automatically, without any human labeling. To this end, we introduce LiDAR-SAM2, a framework that turns a 2D video foundation model, SAM2, into a scalable source of supervision for the 4D LiDAR domain. On the data side, it automatically generates temporally coherent LiDAR-level labels from SAM2 video masks through multi-view projection and spatio-temporal aggregation. On the modeling side, a tailored modality interface and a two-stage learning objective adapt SAM2's video segmentation kernel to spatio-temporal LiDAR structure, so that a single click per object yields a consistent mask track across the sequence. Trained with no human LiDAR annotation, LiDAR-SAM2 produces semantic and panoptic labels on SemanticKITTI that approach the quality of full human annotation from only a few points, and models trained on these labels approach the performance of full ground-truth supervision. This positions LiDAR-SAM2 as a scalable labeling tool that substantially reduces the annotation burden for 3D and 4D scene understanding.
Jihun Kim, Hyun-Kurl Jang, Hyemin Yang et al.· 0 citations
In this study, we propose a method for automatically detecting hard examples from unlabeled real-world data that lead to performance degradation in semantic segmentation models. The proposed method estimates obstacle regions using geometric information from a 3D LiDAR and generates obstacle masks by projecting them onto the image plane. These masks are compared pixel-wise with the semantic segmentation results, and a score based on the misclassification rate within obstacle regions is computed for each image. For evaluation, a semantic segmentation model trained on data collected at the Ikuta Campus of Meiji University was applied to unseen real-world data collected on the Tsukuba Challenge 2025 verification course. Experimental results confirmed that the proposed method could effectively identify images with many misclassified obstacle pixels.
LiDAR semantic segmentation is significant in applications such as autonomous driving and robot navigation, as it greatly improves scene perception and object detection. However, the existing methods face the challenges of achieving high segmentation accuracy while maintaining low computational cost and complexity. In this paper, we propose a new, to our knowledge, efficient and accurate semantic segmentation network for LiDAR, called Range-FDSeg. To reduce the risk of information compression and loss when projecting 3D point cloud data onto 2D range images, we design a multi-channel fusion interactive learning (FIL) module. This module effectively integrates multimodal channels, such as coordinates, depth, and reflectivity, for interactive learning. As a result, FIL module can reduce the noise interference inherent in individual channels and capture the underlying relationships between different physical quantities. To further improve the performance, we introduce a lightweight and dynamic upsampler, called Dysample-S+. It effectively resolves the inherent challenges of traditional sampling methods through its adaptive weighting mechanism, which dynamically adjusts to local geometric patterns and density variations in raw point clouds. Extensive evaluations on publicly available benchmark datasets, including SemanticKITTI, SemanticPOSS, and NuScenes, demonstrate that the proposed Range-FDSeg outperforms most existing state-of-the-art methods.
Hong Tan, Fen Chen, Tingna Liu et al.· Applied Optics· 0 citations
Abstract. Automatic updating of topographic maps remains a significant challenge, as current workflows still rely heavily on manual interpretation of airborne data. This study proposes a method for identifying topographic changes by learning object representations from existing maps and using them as reference data for change detection. Map-derived labels are used to train independent 2D and 3D segmentation networks that generate semantic predictions from orthoimages and point clouds. Unlike conventional change-detection approaches that require temporally aligned datasets of the same modality, the proposed method directly compares newly acquired airborne data with existing map vectors. Semantic predictions from both modalities are vectorized and selectively fused into polygon geometries, which are subsequently compared with reference map vectors to identify object-level “from–to” changes. The workflow highlights potential change regions and their predicted semantic classes, allowing operators to focus inspection on relevant areas rather than the entire dataset. Detected changes include both real-world developments, such as new construction and demolitions, and inconsistencies in the reference map caused by outdated or inaccurate delineations. To assess the effect of multimodal integration, the workflow is compared with a 2D-only baseline. The results indicate that integrating 3D geometric information can reduce noisy detections and improve the spatial consistency of candidate change objects, particularly for water and bridge classes.
G. Anjanappa, Sander J. Oude Elberink· ISPRS Annals of the Photogra...· 0 citations