A more flexible framework in which a predictive model determines the nominal distribution and a separate model estimates a data-dependent radius is developed, which treats calibration as a practical mechanism for reliable decision making rather than a universal guarantee of improved optimization performance.
Abstract
Wasserstein distributionally robust optimization (DRO) is commonly built around the empirical distribution, with the ambiguity radius selected from a concentration bound. Although this construction provides useful statistical guarantees, it can be conservative and does not fully exploit predictive information about the underlying distribution or the difficulty of a particular decision problem. We develop a more flexible framework in which a predictive model determines the nominal distribution and a separate model estimates a data-dependent radius. The key requirement is not that the ambiguity set be centered at the empirical distribution, but that it contain the unknown data-generating distribution with the desired probability. We establish finite-sample guarantees and asymptotic consistency for arbitrary learned centers, derive tractable reformulations for non-uniform discrete predictive distributions, separate predictive-model and scenario-discretization errors, and prove stability under simultaneous perturbations of the center and radius. We further characterize the oracle conditional-quantile radius as the smallest conditionally valid rule and introduce a split-conformal procedure for finite-sample marginal calibration. Experiments on newsvendor problems, synthetic portfolios, distribution shifts, and real financial data show that learned and calibrated ambiguity sets can improve reliability, but do not automatically yield smaller radii or better decisions. Overall, the proposed framework treats calibration as a practical mechanism for reliable decision making rather than a universal guarantee of improved optimization performance.
Distributionally robust optimization (DRO) provides a principled framework for decision-making under distributional uncertainty. Classical data-driven DRO frameworks typically construct ambiguity sets from distributional information, such as moment constraints, divergence neighborhoods, or Wasserstein balls, specified before the downstream loss is considered. We propose a task-aware DRO framework based on targeted integral probability metrics. The ambiguity set is defined directly through the loss functions induced by feasible decisions, thereby controlling the loss discrepancy between an adversarial distribution and a data-driven reference distribution. This construction leads to an expected hinge-constrained formulation that is equivalent to an infinitely constrained loss-discrepancy formulation. It also yields finite-sample guarantees that bypass the ambient curse of dimensionality: whenever an appropriate scalar pointwise concentration inequality is available for the induced loss estimator, the ambiguity radius can be calibrated at the canonical $\widetilde{\mathcal O}(N^{-1/2})$ rate after uniformization over the decision class. As a result, the framework applies broadly to settings including heavier-tailed sub-Weibull losses, Markovian data, outlier-corrupted data, and incomplete data. We derive exact infinite-dimensional dual reformulations, establish out-of-sample and excess-risk guarantees, and develop a conservative Monte Carlo approximation scheme with convergence and suboptimality guarantees. For piecewise affine losses, the sampled problems admit tractable conic reformulations. Numerical experiments in inventory management under heavy-tailed demand and regression with outlier corruption demonstrate strong out-of-sample performance relative to existing approaches.
L. Fang, Jianqiang Cheng, G. A. Hanasusanto et al.· 0 citations
A novel formulation of a Wasserstein-2 metric that uses the Bures-Wasserstein (BW) metric over probability measures with finite second moments is developed, which allows the worst-case distribution to endogenously determine both how many mixture components receive mass and where their means and covariances lie within a continuous support.
Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. The core insight is to select a subset of samples whose empirical distribution maximizes its Wasserstein distance to the fully contaminated empirical distribution, thereby preferentially isolating and removing geometrically influential outliers. To render this optimization computationally tractable, we introduce three algorithms: a marginal screening scheme, SinkMarg, and two joint optimization algorithms, SinkWF and SlicedWF, leveraging entropic optimal transport and sliced Wasserstein approximations, respectively. On the theoretical front, we introduce the Far Exclusion and Local Projection (FELP) contamination model, which characterizes corruptions consisting of well-separated outliers and locally indistinguishable perturbations. Under this model, we prove that the WF estimator achieves minimax optimality over distribution families with bounded covariance. Extensive numerical experiments on synthetic datasets, benchmark anomaly detection suites, and robust generative learning with diffusion models demonstrate that WF serves as a highly practical, model-agnostic preprocessing tool. It delivers competitive outlier detection performance and provides substantial downstream benefits for generative modeling under heavy contamination.
Predict-then-optimize systems usually compress uncertainty into a point forecast and then solve a downstream optimization problem as if the forecast were reliable. Distributionally robust optimization (DRO) offers protection against misspecification, but the ambiguity set is often centered at historical samples and uses a fixed radius. We propose \emph{learned predictive ambiguity sets} (LPAS): a deep contextual model outputs a finite nominal scenario distribution, a state-dependent Wasserstein radius, and optionally an anisotropic ground metric. These outputs define a contextual ambiguity set that feeds a DRO decision layer. The radius is trained by a combination of conditional quantile calibration, size regularization, and downstream decision loss, so that robustness is adaptive rather than globally fixed. We derive the finite dual form used by the decision layer, present a staged training algorithm, and evaluate the method on distributionally robust portfolio optimization with 20 S&P 500 constituents from 2018--2026. The proposed method substantially improves over equal-weight, predict-then-optimize, and historical Wasserstein DRO baselines, achieving 26.28% annualized return, Sharpe ratio 1.30, final wealth 1.61, and lower tail loss than a deep fixed-radius DRO baseline while using a smaller average radius. The results show that learned ambiguity radii can recover most of the performance of strong fixed-radius DRO while reducing unnecessary conservatism and improving regime adaptivity.
It is shown that the classical Bayesian bootstrap closes this gap in U-calibration, which asks one online probability fore-caster to have low regret for every bounded proper loss, including losses unknown when the forecasts are made.
A novel perturbation test based on a nonsmooth max-difference revenue statistic comparing the best null assortment with the best alternative assortment and asymptotic validity of the proposed p-value under adaptive assortment selection is proposed.