A trajectory-based diagnostic framework built on a topology-preserving state space for visually and quantitatively comparing forecast behavior across different AI weather forecast models and shows that the proposed representation provides an intuitive way to compare model-dependent forecast evolution, identify regions of relatively good or poor trajectory behavior, and examine state-or season-dependent differences that are not readily summarized by variable-wise RMSE curves alone.
This study presents OCELOT (Observation-Centric Estimation and Learning for Outlook Trajectories), a global machine-learning forecasting system that predicts future Earth observations directly from heterogeneous satellite and in-situ measurements. Unlike data-driven weather models trained on gridded reanalysis states, OCELOT operates natively in observation space, preserving instrument-specific sampling, viewing geometry, and measurement characteristics. The system combines per-instrument graph-attention encoders, a shared spherical icosahedral latent mesh, a hybrid sliding-window Transformer/spatial graph neural network processor, and metadata-conditioned decoders to produce forecasts up to 12 h ahead. OCELOT is trained on observations for the years 2015 through 2023, validated on the year 2024, and evaluated out of sample on 2025 observations across satellite radiances, radiosondes, aircraft, and surface networks. In the 2025 evaluation, OCELOT produces spatially coherent +12 h forecasts across independent observing systems: microwave temperature-sounding channels show RMSE values of 1.24-1.87 K, while the more surface- and cloud-sensitive AVHRR infrared window channel shows a higher RMSE of 3.95 K. Vertical profile diagnostics show physically consistent radiosonde and aircraft temperature structure. Surface forecasts remain stable through 12 h, with 2-m air-temperature RMSE increasing from about 3.2 K at +3 h to about 3.6 K at +12 h. In paired observation-space comparisons, OCELOT remains less accurate than operational GFS but substantially outperforms persistence at longer lead times for 2-m temperature and 10-m wind components. These results demonstrate that observation-space forecasting can recover large-scale atmospheric structure and provide meaningful short-range skill without reanalysis supervision.
A. Gholoubi, R. McLaren, Mu-Chieh Ko et al.· 0 citations
Many forecast applications require high frequency temporal resolution, yet most state-of-the-art data-driven weather forecasting systems operate at 6-hourly resolution. Although direct hourly forecasting is possible, it suffers from error accumulation and temporal inconsistency. We introduce HourGlass, a probabilistic data-driven temporal downscaling method that reconstructs the evolution between forecast states. HourGlass is trained using variants of the continuous ranked probability score (CRPS) preserving small-scale spatial variability while encouraging temporal consistency. Unlike existing deterministic temporal downscaling approaches, which tend to produce overly smooth fields, HourGlass generates realistic probabilistic forecasts. Training on forecast trajectories rather than reanalysis or analysis data also avoids the temporal inconsistencies present in datasets used by previous methods. We evaluate HourGlass in two settings: AIFS-HourGlass, applied globally to ECMWF's AIFS-Single and AIFS-ENS forecast systems, and Bris-HourGlass, applied regionally to MET Norway's high-resolution stretched-grid ensemble model, Bris. Verification against observations shows that both models retain the skill of their underlying forecasting systems while producing temporally coherent hourly forecasts with realistic small-scale variability. Case studies demonstrate physically consistent evolution during rapidly developing weather events, including extratropical cyclones and organised convection. Hourly precipitation remains challenging: HourGlass improves the spatial realism of precipitation fields but still underestimates the most intense extremes, a common limitation of data-driven weather forecasting models. These results demonstrate that HourGlass effectively bridges the gap between 6-hourly data-driven forecasts and the hourly products required for operational regional and global forecasting.
Magnus Sikora Ingstad, Mariana Clare, Olav Ersland et al.· 2 citations
We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weather forecasting model to Mars. While GraphCast achieves state-of-the-art performance for terrestrial forecasting, its applicability to non-Earth environments remains unexplored. Using the Mars Climate Database (MCD), which provides global atmospheric fields across vertical altitude levels (similar to Earth pressure levels), we evaluate zero-shot and fine-tuned GraphCast predictions of Martian temperature and wind fields. Zero-shot forecasts produce a surprisingly accurate depiction of current conditions but fail to reproduce diurnal variability and rapidly decay toward climatological mean states. To address this limitation, we fine-tune GraphCast using MCD variables and top-of-atmosphere solar radiation forcing while holding humidity constant. Fine-tuning enables rapid learning of Martian thermal variability. Within as few as 10 training epochs, the model begins to capture the diurnal cycle and forecasts up to 10 days reproduce seasonal and vertical temperature structure. Prediction quality improves with training sample size and exhibits sensitivity to seasonal initialization. These results demonstrate that Earth-trained AI weather models can be adapted to simulate Martian atmospheric dynamics, providing a pathway toward rapid planetary weather prediction to support mission operations, dust storm risk mitigation, and future human exploration.
M. Carroll, J. Li, S. Guzewich et al.· 0 citations
Weather forecasting and climate projection frequently use multi-model ensembles (MMEs) to improve short-term forecasts by averaging across models. However, this practice is often not well justified or validated. Using reservoir computing (RC) as a computationally efficient alternative to large-scale physical models, we assess the validity of the MME approach for chaotic time series. By training multiple randomly constructed RCs on the same dataset, we create a multi-model ensemble in which each model has its own unique error. These model errors lead to very different forecasting performances, with forecast error distributions that exhibit heavy tails. The arithmetic mean across forecasts from multiple models for the same target is usually closer to the ground truth than most individual forecasts, and further improvement is achieved by weighted arithmetic means where the weights are constructed based on each model's test-set performance. We show that iterated forecasts over many time steps deviate from the ground truth along the unstable manifold of the target point, in both directions, so that, if forecast errors were independent and had zero mean, the arithmetic mean forecast should approach the true target like $1/\sqrt{\nens}$ where $\nens$ is the size of the multi-model ensemble. We observe deviations from this behavior, which we attribute to the tails of the error distribution of random RCs.
Daniel Estevez Moya, Francesco Martinuzzi, Edmilson Roque dos Santos et al.· 0 citations
Global weather forecast models are vital tools with numerous applications, including public safety, agriculture, and transportation. Recent advancements in artificial intelligence (AI) and deep learning (DL) have shown the potential to enhance weather forecasting accuracy and speed. In this study, we developed a short‐term hourly weather forecast framework with a wavelet transform function for data preprocessing and a spatiotemporal DL model, PredRNN, for predicting five surface atmospheric variables, including wind speed and direction, mean sea level pressure (MSLP), temperature, and precipitation. The framework demonstrated promising results. It produces global forecasts at 0.25° (∼25 km) with a 1‐day lead time RMSE of 1.8 m/s for wind components, 180 Pa for MSLP, and 1.8 K for temperature. Although our model does not surpass state‐of‐the‐art AI weather forecast models across all metrics, it outperforms these models in precipitation forecasting and wind prediction at short lead times and achieves comparable accuracy for MSLP. Its native hourly forecasting capability, together with training on widely accessible GPU hardware, contributes meaningfully to the advancement of accessible DL weather forecasting methods. Our work highlights the importance of integrating temporal components and data transformation techniques to improve the predictability and accuracy of weather forecasts.
Hoang Tran, Hao Li, Vinh Ngoc Tran et al.· Journal of Geophysical Resea...· 0 citations
Machine learning-based weather prediction is revolutionizing weather forecasting by learning from weather data in present-day climate. However, generalization to other climates remains a major challenge. With melting sea ice, land-use change, and increasing ocean temperatures, boundary conditions are changing. Therefore, generalization in time depends on generalization in space. Here, we present three test cases to evaluate whether machine learning-based weather and climate models generalize in space and apply them to GraphCast and NeuralGCM. We reverse or rotate the planet in longitude or latitude under the model's coordinate system and adapt all boundary conditions and forcings accordingly. Physics-based general circulation models simulate a rotated/reversed planet with only rounding errors, but GraphCast and NeuralGCM fail these tests. The analyses furthermore revealed unphysical variable mappings based on correlation rather than causation. We argue that machine learning-based climate models should be designed to pass generalization tests to prevent overfitting on present-day regional climate.
Maren Höver, Milan Klöwer, Christian Schroeder de Witt et al.· 0 citations