Skip to content

Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts

Aug 2026 · 0 citations
Computer Science

TL;DR

The proposed Dynamic Regime-Aware Conformal Prediction (DRACP), which combines density-ratio, localized kernel and probabilistic regime-aware weighting with a self-tuning online significance controller in a unified weighted conformal calibration framework, provides the most reliable calibration.

Abstract

Conformal prediction provides distribution-free prediction intervals but relies on exchangeability, an assumption often violated in economic forecasting because of covariate shift, concept drift, local heterogeneity and latent regimes. We propose Dynamic Regime-Aware Conformal Prediction (DRACP), which combines density-ratio, localized kernel and probabilistic regime-aware weighting with a self-tuning online significance controller in a unified weighted conformal calibration framework. We distinguish three theoretical results: finite-sample validity under oracle importance weights, a coverage-gap bound for estimated weights with rates in effective sample size, and deterministic or regret guarantees for the online controller. We evaluate DRACP against six baselines on 48 real forecasting series covering euro-area and EU-27 HICP inflation, US macroeconomic and energy indicators, and daily financial series. Recent online methods (FACI, strongly-adaptive online conformal prediction and conformal PID) were verified against the authors'implementations. DRACP is not the most efficient method: strongly-adaptive online conformal prediction achieves the best interval score and intervals about 20% narrower. Instead, DRACP provides the most reliable calibration, achieving coverage closest to the nominal 0.90 (0.890), never falling below 0.80 on any series, maintaining the best coverage at all forecast horizons, and performing best during the 2021-2023 inflation surge. The strongly-adaptive method undercovers on 20 of 48 series versus 10 for DRACP. DRACP therefore offers a principled trade-off between calibration and efficiency, favoring reliable coverage when prediction intervals must satisfy coverage standards. An ablation study shows that the online controller and conditional-scale normalization provide most of the performance gain, whereas the weighting components make a smaller contribution.

View source

Similar papers

Open access Jul 2026

Graph-Based Uncertainty-Aware Financial Forecasting via Cross-Asset Conformal Prediction

Reliable financial forecasting requires not only accurate point predictions but calibrated uncertainty that remains valid under market stress. Conformal prediction offers distribution-free, finite-sample coverage guarantees and has recently been applied to risk-adjusted financial models, yet existing approaches calibrate each asset in isolation. Per-asset calibration is statistically inefficient, degrades sharply when historical data are scarce, and yields poor conditional coverage precisely when it matters most, that is, during volatile and highly correlated market regimes. We propose Cross-Asset Graph Conformal Prediction (CA-GCP), a framework that pools volatility-normalized nonconformity scores across the correlation-graph neighborhood of each target asset using a proximity- and recency-weighted quantile, grounded in the theory of weighted conformal prediction. A lightweight systemic-stress modulator further widens intervals on days of market-wide turbulence. On five years of daily returns for 100 S&P 500 constituents, CA-GCP reduces the cross-sectional standard deviation of per-asset coverage from 1.55% to 0.96% and improves worst-decile coverage from 91.4% to 94.1% relative to a faithful re-implementation of a state-of-the-art per-asset volatility-adaptive conformal baseline, while achieving 95.2% coverage on extreme-volatility days versus 90.4%. Under severe calibration scarcity, with as few as 20 samples per asset, CA-GCP maintains a coverage standard deviation below 0.9%, four to five times more stable than per-asset methods. The gains are robust across graph topologies and forecasting backbones, indicating that cross-asset pooling, rather than any particular graph, is the source of improvement. CA-GCP is model-agnostic, adds negligible computational overhead, and comes with finite-sample validity bounds.

Ethan Parker, Rui Zhang · 0 citations
Open access Jul 2026

Regime-Aware Conformal Transformer for Multi-Horizon Financial Forecasting Under Market Uncertainty

Financial forecasting systems deployed in portfolio monitoring and risk control require not only point forecasts, but also prediction intervals that remain meaningful when markets change volatility regimes. This paper proposes a Regime-Aware Conformal Transformer (RACT) for multi-horizon financial forecasting under market uncertainty. The model combines a compact Transformer encoder with asset and volatility-regime tokens, producing simultaneous return forecasts for 1-, 2-, 4-, and 8-week horizons. To quantify uncertainty, we introduce a regime-aware conformal calibration rule that computes horizon-specific volatility-normalized residual quantiles inside low-, medium-, and high-volatility regimes, while retaining a global conformal safety floor to avoid under-coverage when local volatility scaling becomes too optimistic. The resulting interval construction is conservative, causally ordered, and easy to reproduce. Experiments use the real Plotly Express stocks dataset containing weekly normalized closing prices for six technology stocks in 2018/2019. On this compact public dataset, RACT with regime-aware conformal calibration achieves 98.3% average empirical coverage at a 90% nominal level and improves 8-week coverage from 76.5% for pure volatility-scaled conformal calibration to 94.6%, at the cost of wider intervals. The results show that regime-aware conformal envelopes can materially improve long-horizon reliability, although they should be validated on larger trading datasets before production deployment.

J. Stein · 0 citations
Open access Sep 2026

Reliable value at risk estimation with conformal prediction

Value-at-Risk (VaR), the most widely used measure of market risk, is typically evaluated through backtesting of point forecasts. Such procedures, however, say little about the uncertainty of the estimated quantile. Existing interval methods are each tied to a specific model class and fail when its underlying assumptions are violated. We propose Quantile Dynamically-Tuned Adaptive Conformal Inference (QDtACI), a model-agnostic conformal calibration layer that constructs finite-sample intervals around any VaR forecast, using only the return series and the forecast itself. QDtACI adapts dynamically-tuned adaptive conformal inference to the quantile setting through two components: a pinball-loss nonconformity score aligned with the quantile objective and an asymmetric interval construction, and a multi-speed expert-aggregation mechanism driven by a smoothed violation error and a composite loss on coverage, width, and stability. On synthetic GARCH data, where the true VaR is observable, QDtACI attains near-nominal coverage of the true VaR when the underlying forecast is well-specified, and its coverage degrades in a controlled way as the forecast is misspecified. Against the Delta method, a bootstrap, and the DtACI baseline, it achieves coverage closer to nominal at comparable or better interval quality (Winkler score). Applied to a portfolio of 24 fixed-income assets (2016–2024) with VaR forecasts from CAViaR, DCC-GARCH, and copula models, the intervals are stable in calm periods and widen sharply during stress, including the COVID-19 shock and the 2022–2023 monetary tightening. Because true coverage cannot be measured on real data, we further provide a return-only diagnostic that indicates when interval calibration can be trusted as a proxy for coverage of the true VaR. QDtACI thus offers a single, broadly applicable procedure for uncertainty quantification in VaR, whose reliability tracks the quality of the underlying forecast.

Milo Ivancevic, K. Nguyen, Zhiyuan Luo · 0 citations
#artificial intelligence Preprint Aug 2026

SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting

SPACE is proposed, a conformal wrapper for sample-generating multivariate forecasters that consistently brings realized joint and rolling coverage closer to the nominal target, achieving superior coverage-efficiency tradeoffs relative to competing wrappers.

Baishi Li, Kelvin J. L. Koa, Ke-Wei Huang · 0 citations
Preprint Aug 2026

Conformal Kelly: Conformal Prediction Intervals as the Scale in Fractional Kelly Position Sizing

Conformal prediction has traditionally been used to quantify prediction uncertainty. We put that uncertainty to a second use, combining a 75% conformal interval with fractional Kelly to size portfolio positions: as the range widens we shrink the position, and as it narrows we grow it. On a six-year development window (2016-2021), with trading costs and strict leverage caps, this compounds at 28.5% annualised net log growth with a Sharpe ratio of 1.34 and a 27.7% maximum drawdown, versus 15.9% for holding the S&P 500 and 21-22% for passive portfolios at the same leverage. Our main development-window finding runs against the literature's advice for conformal prediction on time series. Every tweak that adapts the interval faster to market conditions costs 0.7 to 5.3 points of annual growth; the winner is the simplest method: slow, unweighted, per-asset rolling quantiles. When an interval sizes a position rather than describing one forecast, width stability beats local sharpness. It also beats the textbook standard deviation by 2.1 points at matched leverage. We also implement a risk control: when the intervals miss on the downside far more than their historical rate, we cut leverage. On the development window this cut maximum drawdown from 27.7% to 20.3% while raising the Sharpe ratio, beating all 40 placebo timings (rank-based p = 1/41). These numbers came from an autonomous LLM-agent search over 200 configurations, so we sealed all data from 2022 onward and pre-registered configurations, benchmarks, and interpretation rules before one evaluation. Calibration held (0.745 coverage against 0.750, weakest through 2022); growth did not: the two configurations earned 8.5% and 7.0% per year, below the passive benchmarks, and a pre-registered hindsight benchmark beat them on raw growth while taking a 46% drawdown. All outcomes are reported as pre-registered.

Robert Jacob Ryan · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.