Uncertainty-aware large language models: a scoping review of conformal prediction methods.
In recent years, there has been growing interest in applying conformal prediction (CP) to large language models (LLMs) across diverse domains to enhance the trustworthiness of their predictions. However, the literature still lacks a comprehensive survey of this rapidly emerging area. Thus, to fill this gap, this article presents a comprehensive review and analysis of CP for LLMs. We review over 106 studies and propose a novel taxonomy that categorises existing methods into six groups. In addition, we observe trends in the literature with regard to LLM, dataset, and task selection, and make recommendations for researchers to improve task diversity and address gaps in black-box LLM uncertainty estimation. Interestingly, we find that logit-free methods tend to outperform logit-based methods, with a lower mean absolute coverage error and prediction set size. This suggests that logit-free methods, which rely on uncertainty signals based on self-consistency sampling, semantic diversity, and other methods, may sidestep known issues with miscalibrated token probabilities and may have advantageous robustness to tokenisation and decoding idiosyncrasies, particularly for open-ended generation where the performance gap is more pronounced. This indicates that black-box uncertainty signals may more directly capture semantic correctness or answer stability. We also find that on average, conformal methods for large vision language models (LVLMs) have higher overcoverage error than LLMs, and almost non-existent undercoverage error, suggesting that methods for LVLMs may be more conservative. While we do not claim a definitive causal explanation, empirical evidence suggests that conformal methods for LVLMs exhibit a stronger coverage-informativeness trade-off than those for LLMs. This article is part of the discussion meeting issue 'Advancing uncertainty quantification in AI systems'.