Skip to content
Preprint

AlignJEPA: Predictive Vision-Language Alignment for Remote Sensing Foundation Models

Aug 2026 · 0 citations · 45 references
Computer Science

TL;DR

AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language by using a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network.

Abstract

Remote sensing (RS) foundation models provide transferable Earth observation representations across sensors, resolutions, and geographies, yet most remain weakly aligned with natural language, limiting natural-language archive search, image-text retrieval, and question-conditioned analysis. We propose AlignJEPA, a JEPA-inspired predictive vision-language alignment framework for remote sensing foundation models. AlignJEPA uses a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network. Instead of relying on global image--text contrastive alignment alone, the framework predicts remote-sensing text embeddings from masked visual foundation-model tokens. Its mask-aware multi-scale predictive aligner aggregates visible tokens at fine, regional, and global scales, jointly models them with a cross-scale Transformer, and projects the resulting representation into the text space using learned query pooling. Training combines semantic prediction with bidirectional contrastive retrieval. We train and evaluate AlignJEPA on BigEarthNet.txt for natural-language Sentinel retrieval, evaluate cross-dataset adaptation on RSICD, and use RSVQA only as a closed-set representation probe. AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language.

View source

Similar papers

Preprint Aug 2026

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short seque...

Yupan Ding, Jing Xiao, Zhenyuan Zhang et al. · 0 citations
May 2024

PriorCLIP: Visual-Prior-Guided Vision–Language Model for Remote Sensing Image–Text Retrieval

PriorCLIP is a visual-prior-guided vision–language model that organizes closed-domain and open-domain retrieval as two complementary realizations of the same learning principle, which introduces remote sensing scene knowledge as a visual prior, then uses this prior to adapt image and text representations to the availab...

Jiancheng Pan, Muyuan Ma, Qing Ma et al. · 12 citations · ⚡1
Open access Aug 2026

GeoGATE: Geo-Sensor-Guided Adaptive Token and Evidence Reasoning for High-Resolution Remote Sensing Image Understanding

GeoGATE is introduced, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning that associates adaptive slicing most strongly with localization, retrieval with language and...

Jing-Nan Zhang, Feng-Jun Zhang · 0 citations
2026

MMF-Net: Multimodal Mutual Feedback Fusion Network for Referring Remote Sensing Image Segmentation

Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations shou...

Chongyang Li, Chen Wang, Wen-Kai Zhang et al. · 0 citations
Conference Open access Sep 2026

Unified Sequence Modeling for Remote Sensing: A Parameter-Efficient Foundation Model via Prompt-Driven Granularity Alignment

RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.

Yang Liu, Wei-Xing Luo, Huai-Zhou Qi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.