AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language by using a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network.
Abstract
Remote sensing (RS) foundation models provide transferable Earth observation representations across sensors, resolutions, and geographies, yet most remain weakly aligned with natural language, limiting natural-language archive search, image-text retrieval, and question-conditioned analysis. We propose AlignJEPA, a JEPA-inspired predictive vision-language alignment framework for remote sensing foundation models. AlignJEPA uses a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network. Instead of relying on global image--text contrastive alignment alone, the framework predicts remote-sensing text embeddings from masked visual foundation-model tokens. Its mask-aware multi-scale predictive aligner aggregates visible tokens at fine, regional, and global scales, jointly models them with a cross-scale Transformer, and projects the resulting representation into the text space using learned query pooling. Training combines semantic prediction with bidirectional contrastive retrieval. We train and evaluate AlignJEPA on BigEarthNet.txt for natural-language Sentinel retrieval, evaluate cross-dataset adaptation on RSICD, and use RSVQA only as a closed-set representation probe. AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language.
Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short seque...
Yupan Ding, Jing Xiao, Zhenyuan Zhang et al.· 0 citations
DARAD is proposed, a dual-adapter and ranking-aware distillation framework that preserves the historical cross-modal ranking structure while learning new visual and textual concepts from evolving archives to address the challenge of continual RS-ITR.
Xi Chen, Xu Chen, Xiang-Yang Jia et al.· 0 citations
PriorCLIP is a visual-prior-guided vision–language model that organizes closed-domain and open-domain retrieval as two complementary realizations of the same learning principle, which introduces remote sensing scene knowledge as a visual prior, then uses this prior to adapt image and text representations to the availab...
Jiancheng Pan, Muyuan Ma, Qing Ma et al.· IEEE Transactions on Geoscie...· 12 citations· ⚡1
GeoGATE is introduced, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning that associates adaptive slicing most strongly with localization, retrieval with language and...
Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations shou...
Chongyang Li, Chen Wang, Wen-Kai Zhang et al.· IEEE Geoscience and Remote S...· 0 citations
RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.
Yang Liu, Wei-Xing Luo, Huai-Zhou Qi et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.