Skip to content
Preprint

Decision Making Needs Uncertainty Quantification [Lecture Notes]

Jul 2026 · 3 citations · 13 references
Computer Science Mathematics

TL;DR

This lecture note develops, from first principles and within a single decision-theoretic setting, the link between the decision objective and the knowledge of an agent and the form of uncertainty representation that is sufficient to act optimally.

Abstract

Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision maker, or agent, is uncertain, the way that uncertainty is represented decides how well the agent performs and how much its performance can be trusted. This lecture note develops, from first principles and within a single decision-theoretic setting, the link between the {objective} and the knowledge of an agent and the form of uncertainty representation that is sufficient to act optimally. To start, assuming a known environment distribution, we show that a risk-neutral agent needs the posterior distribution over the state, whereas a risk-averse agent can rely without loss of optimality on a {prediction set} and a worst-case decision rule. We then turn to the case in which the environment is unknown, and identify three complementary approaches to address the resulting epistemic uncertainty: calibration of a fixed predictor, credal (ambiguity) sets with distributionally robust optimization, and Bayesian inference over model parameters. The common thread is that reliable decisions require an uncertainty representation matched to the decision objective and to the knowledge profile of the agent, together with a guarantee that certifies the utility the agent will actually obtain.

View source

Similar papers

Preprint Aug 2026

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.

D. Ganguly, Jan Křetinský · 0 citations
Open access Aug 2026

A Bayesian composite risk approach for stochastic optimal control and Markov decision processes

Inspired by Shapiro et al. [74], we consider a stochastic optimal control (SOC) and Markov decision process (MDP) under simultaneous epistemic and aleatoric uncertainties using Bayesian composite risk (BCR) measures. The proposed BCR-SOC/MDP model evaluates the risk of stagewise cost via a two-layer framework: the inner risk measure tackles aleatoric uncertainty conditional on a latent environment parameter, while the outer risk measure deals with the epistemic uncertainty of the inner risk under the Bayesian posterior. The resulting time-varying risk evaluation induced by Bayesian updating enables an information-adaptive risk-sensitive decision framework. Unlike [74], our policies are allowed to depend explicitly on the posterior belief, reflecting that accumulated information about epistemic uncertainty can influence the assessment of future aleatoric uncertainty and, consequently, the decision maker’s actions [79]. The new modeling paradigm subsumes several classical SOC/MDP formulations, including risk-averse and distributionally robust SOC/MDPs as well as partially observed and Bayes-adaptive MDPs, and generates so-called preference robust SOC/MDP models. Moreover, we derive conditions under which the BCR-SOC/MDP model is well-defined, show that finite-horizon BCR-SOC/MDP models can be solved via dynamic programming, and extend the analysis to the infinite-horizon case. Under standard conditions, we establish asymptotic convergence of the optimal values and optimal policies as data accumulate, and provide quantitative error bounds for several representative classes of risk measures. To enhance computational tractability, we develop a hyper-parameter discretization approach for the posterior belief space. Finally, we carry out numerical tests on a spread betting problem and an inventory control problem, demonstrating the effectiveness of the proposed model and numerical schemes.

Wentao Ma, Zhi-Ping Chen, Huifu Xu · 0 citations
Preprint Jul 2026

Subjective Risk Decomposition: A New View for Uncertainty Quantification

We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation, but instead consequences, of higher-level modelling decisions. We show how epistemic and aleatoric uncertainty measures can be derived via decomposition of a subjective risk, based on a strictly proper loss. Reverse cross entropy provides a prominent example, where decomposition recovers the classic information-theoretic uncertainty terms. The same approach recovers numerous measures previously proposed across the UQ literature, providing them a common theoretical foundation. This suggests a new approach to UQ: given a modelling scenario and strictly proper loss, the corresponding epistemic and aleatoric terms are induced by the subjective-risk decomposition. We then extend our view to learning theory: we introduce and analyse subjective risk analogues of excess risk, approximation error and estimation error, and identify the connections to UQ. We consider this a first step towards a full learning-theoretic framework for uncertainty quantification.

R. Alamri, Michele Caprio, Gavin Brown · 0 citations
Preprint Aug 2026

Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies

Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions but potentially different transition probabilities and rewards. Optimizing a single policy across all possible MDPs may sacrifice performance, while preparing an individually optimized policy for every MDP may violate operational, regulatory, or interpretability constraints on the number of policies that can be prepared and deployed. We consider settings in which model uncertainty is resolved shortly before execution, allowing the most suitable policy to be selected from a limited set prepared in advance. We introduce $k$-adaptable policy synthesis, which optimizes such a set of $k$ policies under a minimax-regret objective. We prove that the problem is NP-hard and develop KAPS, an exact nested branch-and-bound algorithm with problem-specific bounds and heuristics. KAPS jointly optimizes which MDPs share a policy and the policies themselves. Experiments across various UMDP benchmarks show that the largest reduction in regret consistently occurs when increasing from one to two policies. In the single-policy setting, KAPS is competitive with existing methods in solution quality and proves optimality substantially more often.

Sterre Lutz, D. Vos, M. Spaan et al. · 0 citations
Open access Jul 2026

Epistemic utilities, self-knowledge and Causal Decision Theory

Causal and Epistemic Decision theories differ in their recommendations in a large number of cases, the most famous of which is Newcomb’s Problem. These cases, including the original one, tend to be outlandish and unusual. We show that if one applies decision theory to epistemic utilities measured by a strictly proper scoring rule, the two theories will deviate in their recommendations in cases that any Bayesian agent can easily put themselves in. Moreover, the recommendations of Causal Decision Theory in these cases are implausible: it recommends that we perform the action whose performance gives us no information regarding the hypothesis with respect to which we are maximizing our epistemic utility. We shall see that requiring choices to be ratifiable, or using Barnett’s recent graded ratifiability criterion, does not help either. Nor is it completely clear that Evidential Decision Theory is off the hook.

Alexander R. Pruss · 0 citations
Preprint Aug 2026

On the Complexity of Bayesian Signal Processing

We develop a computational framework for Bayesian decision-making. We show that as long as no action is optimal in every state, Bayes-optimal choice is intractable. This hardness need not arise from large action, state, or signal spaces, nor from a complicated represented utility function: extracting enough information from a hard-to-interpret signal to act optimally can itself be computationally hard. We also characterize tractability across approximation notions and identify their sources of difficulty. Under the probably approximately correct criterion, sample-based Bayesian learning is tractable if and only if the signal support is bounded. Our results provide justifications for bounded rationality, costly Bayesian inference, and sample-based Bayesian learning.

Yi Liu · 0 citations