Skip to content
Preprint

Towards robust multimodal 3D object detection via visual foundation models

Sep 2026 · 0 citations · 53 references
Computer Science

TL;DR

This work introduces SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information, and develops the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical contextual information.

Abstract

Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and environmental changes. To address this problem, we propose RoboDistill, a robust and generalizable multimodal 3D object detection framework that leverages visual foundation models (VFMs), such as the Segment Anything Model (SAM). First, we introduce SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information. Second, we design the AD Feature Pyramid Network (AD-FPN) to refine and upsample SAM features at multiple scales for seamless fusion with LiDAR features. Third, we develop the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical contextual information. Finally, we introduce KD Fusion, in which the pretrained SAM-AD serves as a teacher that distills high-quality visual knowledge into a lightweight point-cloud network, thereby improving robustness under noisy conditions. Extensive experiments across 27 challenging OOD corruption settings show that RoboDistill generally delivers stronger or competitive detection performance and robustness relative to representative state-of-the-art methods. This work bridges the gap between VFMs and 3D object detection and advances robust multimodal perception for real-world autonomous-driving applications.

View source

Similar papers

Oct 2026

Robust Multimodal Gated Fusion for 3-D Object Detection via Alignment and Denoising

Multimodal perception integrating light detection and ranging (LiDAR) and cameras has become a key paradigm for 3-D object detection, as it leverages both geometric structure and semantic information. However, in real-world autonomous driving scenarios, calibration errors, adverse weather, and sensor degradation can in...

Hui-Lin Huang, Yan Bai, Peng-Yuan Wang et al. · 0 citations
Aug 2026

Fadet: a fusion-aware 3D detection network with cascaded feature enhancement for small object detection in autonomous driving

This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.

Chang-Hong Yu, Shaoshi Luo, Wen-Li Shen · 0 citations
Sep 2026

Class-Enhanced Multi-Sampling and Multi-Level Graph Attention for 3-D Object Detection

Three-dimensional object detection is a key component of the perception module in autonomous driving systems. Compared to camera images, LiDAR point clouds provide richer spatial information, such as detailed structural and geometric cues of objects. However, existing 3D object detection methods face two major challeng...

Xiangyang Wu, Ji-Tao Pan, Qing-Long Jiao et al. · 0 citations
2026

EDAFusion: LiDAR-Guided Multilevel Depth Enhancement and Dynamic Scale Attention Fusion for Multimodal 3-D Object Detection

Three-dimensional object detection plays a critical role in intelligent robotics and autonomous driving, where accurate and robust perception remains challenging under multimodal fusion settings. Existing bird’s-eye-view (BEV)-based multimodal methods still suffer from unreliable camera depth estimation, insufficient a...

Jian-Qiang Su, Ji-Yuan Wang, Yong-Sheng Qi et al. · 0 citations
Open access 2026

Feature Alignment for NeRF-Based 3-D Object Detection

Two lightweight and complementary modules to enhance voxel feature quality for NeRF-based 3D detection with consistent improvements over the NeRF-RPN baseline in both recall and precision are introduced.

Yu-Han Wang, Gang Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.