Skip to content
Open access

DenseRS-CLIP: Enhancing Dense Feature Representation of Remote Sensing CLIP via Attention-Decoupled Dual-Branch Distillation

2026 · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · Vol 19, pp. 29136-29152 · 0 citations · 44 references

Abstract

Existing remote sensing vision–language foundation models mainly follow the CLIP-style global image-text alignment paradigm. While effective for image-level semantic understanding, this paradigm leaves patch-level representations insufficiently discriminative and spatially inconsistent for dense prediction tasks. To address this issue, we propose DenseRS-CLIP, a region-structured two-stage adaptation framework for enhancing dense representations of remote sensing CLIP models. In the first stage, we perform dual-granularity image-text contrastive pretraining to obtain a domain-adapted CLIP model with robust global semantic alignment. In the second stage, we introduce an attention-decoupled dual-branch distillation framework that reuses existing bounding-box and mask-derived localization annotations to construct ROI-level distillation signals without requiring additional region-text descriptions. Specifically, dense features are decoupled into content and context branches. The content branch is optimized by region-structured semantic distillation with a region correlation constraint, which improves local semantic discriminability and suppresses interregion feature homogenization. The context branch is optimized by DINOv3-guided topology distillation, which aligns patch-level self-similarity structures to improve spatial consistency and boundary awareness. Experiments on seven benchmarks covering region classification, visual grounding, and referring image segmentation show that DenseRS-CLIP achieves consistent improvements over representative remote sensing VLFMs and controlled distillation variants. These results validate the effectiveness of weakly localized, branch-specific distillation for dense remote sensing vision–language representation learning.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.