Value-at-Risk (VaR), the most widely used measure of market risk, is typically evaluated through backtesting of point forecasts. Such procedures, however, say little about the uncertainty of the estimated quantile. Existing interval methods are each tied to a specific model class and fail when its underlying assumptions are violated. We propose Quantile Dynamically-Tuned Adaptive Conformal Inference (QDtACI), a model-agnostic conformal calibration layer that constructs finite-sample intervals around any VaR forecast, using only the return series and the forecast itself. QDtACI adapts dynamically-tuned adaptive conformal inference to the quantile setting through two components: a pinball-loss nonconformity score aligned with the quantile objective and an asymmetric interval construction, and a multi-speed expert-aggregation mechanism driven by a smoothed violation error and a composite loss on coverage, width, and stability. On synthetic GARCH data, where the true VaR is observable, QDtACI attains near-nominal coverage of the true VaR when the underlying forecast is well-specified, and its coverage degrades in a controlled way as the forecast is misspecified. Against the Delta method, a bootstrap, and the DtACI baseline, it achieves coverage closer to nominal at comparable or better interval quality (Winkler score). Applied to a portfolio of 24 fixed-income assets (2016–2024) with VaR forecasts from CAViaR, DCC-GARCH, and copula models, the intervals are stable in calm periods and widen sharply during stress, including the COVID-19 shock and the 2022–2023 monetary tightening. Because true coverage cannot be measured on real data, we further provide a return-only diagnostic that indicates when interval calibration can be trusted as a proxy for coverage of the true VaR. QDtACI thus offers a single, broadly applicable procedure for uncertainty quantification in VaR, whose reliability tracks the quality of the underlying forecast.
Milo Ivancevic, K. Nguyen, Zhiyuan Luo· Risk Management· 0 citations
In recent years, there has been growing interest in applying conformal prediction (CP) to large language models (LLMs) across diverse domains to enhance the trustworthiness of their predictions. However, the literature still lacks a comprehensive survey of this rapidly emerging area. Thus, to fill this gap, this article presents a comprehensive review and analysis of CP for LLMs. We review over 106 studies and propose a novel taxonomy that categorises existing methods into six groups. In addition, we observe trends in the literature with regard to LLM, dataset, and task selection, and make recommendations for researchers to improve task diversity and address gaps in black-box LLM uncertainty estimation. Interestingly, we find that logit-free methods tend to outperform logit-based methods, with a lower mean absolute coverage error and prediction set size. This suggests that logit-free methods, which rely on uncertainty signals based on self-consistency sampling, semantic diversity, and other methods, may sidestep known issues with miscalibrated token probabilities and may have advantageous robustness to tokenisation and decoding idiosyncrasies, particularly for open-ended generation where the performance gap is more pronounced. This indicates that black-box uncertainty signals may more directly capture semantic correctness or answer stability. We also find that on average, conformal methods for large vision language models (LVLMs) have higher overcoverage error than LLMs, and almost non-existent undercoverage error, suggesting that methods for LVLMs may be more conservative. While we do not claim a definitive causal explanation, empirical evidence suggests that conformal methods for LVLMs exhibit a stronger coverage-informativeness trade-off than those for LLMs. This article is part of the discussion meeting issue 'Advancing uncertainty quantification in AI systems'.
Alice E. Ashby, Khuong An Nguyen, Zhiyuan Luo· Philosophical transactions....· 1 citation