Skip to content
Preprint

Post-Learning Inference for Combinatorial Optimizers with High-Dimensional Sparse Contextual Information via Minimal Directional Perturbation

Jul 2026 · 0 citations
Mathematics

TL;DR

A novel perturbation test based on a nonsmooth max-difference revenue statistic comparing the best null assortment with the best alternative assortment and asymptotic validity of the proposed p-value under adaptive assortment selection is proposed.

Abstract

We study post-learning inference for structural properties of data-dependent combinatorial optimizers. The target is whether an oracle optimizer, rather than a latent parameter or smooth functional, belongs to a prescribed class, such as a category-mix, inventory, or resource-feasibility class. We focus on a high-dimensional contextual multinomial logit model with sequentially adaptive data collection, where the parameter-to-optimizer map is discontinuous and the policy induces temporal dependence. We propose a novel perturbation test based on a nonsmooth max-difference revenue statistic comparing the best null assortment with the best alternative assortment. The test perturbs the estimated terminal revenue surface on the selected support: random unit directions capture directional uncertainty, while the minimal perturbation radius captures magnitude uncertainty and yields a p-value. This localizes inference near the null--alternative boundary and avoids uniform error control over the full candidate class. The data are collected by an \(\ell_1\)-penalized online likelihood policy that performs variable selection while controlling regret. Using a new anti-concentration argument for Gaussian maxima differences and martingale Gaussian coupling, we establish uniform estimation rates, effective support recovery, and asymptotic validity of the proposed p-value under adaptive assortment selection. We prove asymptotic size control and power consistency under a localized signal condition.

View source

Similar papers

Preprint Jul 2026

Harnessing Heterogeneous Data for Conditional Optimization via Optimal Transport

Conditional optimization tailors decisions to contextual or event information, but its practical use is often limited by the difficulty of learning the relevant conditional distribution from finite samples of a target joint distribution. This challenge is especially acute when target joint data are scarce or unavailable, or when few observations fall in the conditioning region of interest. Related joint data may be available from multiple sources, such as different stores, markets, populations, or operating environments, but these sources may be biased relative to the target distribution and cannot be pooled naively. We develop a distributionally robust framework based on optimal transport (OT) for harnessing such heterogeneous data in conditional optimization. The framework constructs ambiguity sets over joint distributions using OT distances to empirical source distributions and optimizes worst-case conditional performance over plausible target laws. We propose three OT ambiguity sets that capture different ways of using heterogeneous sources: enforcing simultaneous source consistency, aggregating source discrepancies through weights, and centering the ambiguity set at an OT barycenter. We derive tractable reformulations, establish feasibility conditions, discuss parameter choices, and characterize the relationships among the formulations, revealing trade-offs between robustness, information aggregation, and computational complexity. We demonstrate the value of the framework through a conditional assortment problem using demand and product-feature data from multiple stores.

Jonathan Yu-Meng Li, Qinyu Wu · 0 citations
Preprint Jul 2026

Learning the Center and Radius of Wasserstein Ambiguity Sets for Data-Driven Decision Making

A more flexible framework in which a predictive model determines the nominal distribution and a separate model estimates a data-dependent radius is developed, which treats calibration as a practical mechanism for reliable decision making rather than a universal guarantee of improved optimization performance.

Junjie Guo · 1 citation
Preprint Jul 2026

A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents'risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics. On the elicitation side, we propose an adaptive Bayesian IRL method that infers agents'latent risk objectives from their noisy observed decisions, explicitly allowing agents to take stochastic and suboptimal actions. We establish the existence of a finite set of distinguishing questions that identifies the preferred distortion riskmetric within the candidate class and prove that the convergence rate of the algorithm is of order $O(\exp(-cm+O(\sqrt{m\log m})))$ under general settings, where $c>0$ is a constant and $m$ denotes the number of algorithm iterations. On the optimization side, we develop a model-free RL algorithm for optimizing policies under conditional distortion riskmetrics. By representing the objective as an integral of the conditional cost quantile function with respect to the distortion function, the method unifies distortion-riskmetric objectives. We optimize diverse risk objectives by extending the Proximal Policy Optimization (PPO) algorithm with policy, value, and quantile neural networks, where the quantile network estimates the full conditional cost quantile function and enables numerical evaluation of general risk objectives. A comprehensive empirical study demonstrates the framework's elicitation accuracy and effectiveness in complex financial environments.

Yang Liu, Yuhao Liu, Yunran Wei · 0 citations
Preprint Aug 2026

Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss

Outcomes are regressed on a calibrated probability vector for unobserved class membership. Under a structural conditional mean excluding the score and conditional calibration, the observed-data model reduces to a partially linear regression. The probability vector is a Berkson-type surrogate for membership, so the effect vector $\tau$ is identified without attenuation. In practice the vector is often coarsened to a hard label - an argmax, a confidence threshold - and that label need not retain the Berkson property. For any coarsening the plug-in estimator converges to $\mathcal{A}\tau$, where $\mathcal{A}-I$ is determined by the regression of the discarded part of the score on the retained part. Coarsening therefore leaves $\tau$ undistorted exactly when that regression vanishes, and otherwise distorts some contrasts far more than others. The same operator determines the bias that drives coverage loss. Where that bias is of the order of the standard error, the Wald interval has limiting coverage $\Phi(z-\nu)-\Phi(-z-\nu)$, with $\nu$ their ratio. A fixed bias sends coverage to zero. Operator, standard error, and - through the uncoarsened estimator - the bias are estimable from observed data, so the implied coverage can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening. Three real-data audits exhibit the direction-specific distortion.

M. T. Kurbucz · 0 citations
Preprint Aug 2026

From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation

Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When the business goal is incremental impact, as in marketing campaigns, incentives, and notifications, this paradigm systematically misallocates resources toward users who would have acted anyway. We present a decision-centric framework that instead optimizes causal effects under global constraints, aligning three components under a single objective: a causal neural network with a Transformer backbone for individual treatment-effect estimation, a Bayesian neural-bandit layer for uncertainty-aware exploration, and a dual-based large-scale linear-programming layer for constrained allocation. The framework also supports sequential context and multi-outcome, attribute-conditioned scoring through a Transformer encoder and outcome embeddings. We evaluate it with offline simulations on a public bandit dataset, targeted architectural ablations, and an online A/B test on LinkedIn Feed marketing traffic. We also distill production lessons on causal training-data construction and cost and delivery control, which were critical to successful deployment. The end-to-end treatment policy delivered a statistically significant $+7.20\%$ lift in the primary long-term-value metric, demonstrating the feasibility of production-scale causal optimization under business constraints.

Changshuai Wei, John Bencina, Phuc Nguyen et al. · 0 citations
Preprint Aug 2026

Duality and Error for Predictively Oriented Inference

This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a unique predictive distribution even though the parameter distributions have no weak limit on the original parameter space.

Aurya Javeed, D. Kouri, Teresa Portone et al. · 0 citations