Skip to content
Preprint

CFR without Unbiasedness: Deterministic Guarantees for Persistent Public-Chance Schedules

Aug 2026 · 0 citations · 23 references
Computer Science

TL;DR

A deterministic target-transfer theorem is established for uniform, nonnested additive public cuts that bounds full-cut exploitability by regret on the delivered feedback and a public-debit term that couples prefix coverage discrepancy with motion along the realized strategy path.

Abstract

At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update. Exact evaluation processes the full cut at one strategy profile; persistent partial evaluation processes a fixed without-replacement order across evolving profiles. The latter covers every outcome once per epoch, yet its feedback is generally conditionally biased because earlier batches influence the profiles seen by later batches. We establish a deterministic target-transfer theorem for uniform, nonnested additive public cuts. The theorem bounds full-cut exploitability by regret on the delivered feedback and a public-debit term that couples prefix coverage discrepancy with motion along the realized strategy path. Consecutively balanced schedules consequently converge for additive signed regret matching (RM) and RM+ under predetermined averaging weights, while a fixed RM+ construction proves that the discrepancy--path product is necessary in general. A component-resolved form of the theorem converts an execution trace into a numerical exploitability certificate. On two released heads-up no-limit hold'em turn endgames, persistent order improves substantially over fresh reshuffling despite identical epochwise coverage, and partial coverage wins every registered shallow matched-budget comparison. A depth study locates a crossover between 32 and 64 full-cut outcome budgets, after which complete coverage dominates. These results characterize public-chance width and order as learning variables and provide a deterministic basis for designing and auditing persistent CFR schedules.

View source

Similar papers

Open access Sep 2026

CFR variant and iterate reporting in small imperfect-information games

Empirical comparisons of counterfactual regret minimization (CFR) variants often mix two distinct choices: the update rule used during training and the iterate reported at evaluation time. This note isolates those choices in an exact small-game setting where exploitability is computed by exact best response, so no me...

Behbod Keshavarzi, H. Navidi · 0 citations
#machine learning Review Sep 2026

Tail-Influence Sampling for CVaR Policy Evaluation

Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to es...

P. Bourigault, Xiao-Tong Ji, Matthieu Zimmer et al. · 0 citations
Preprint Sep 2026

Moment Ambiguity and the Limits of Robust Stochastic Optimization

We study fundamental information-theoretic limits of robust stochastic optimization when the distribution is known only through its exact moment sequence. We develop a unified framework that produces families of distinct distributions sharing all moments yet inducing radically different optimal decisions, thereby estab...

Andrés Cristi, Matteo Russo, Jie-Chen Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification

Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budg...

Miao-Bo Hu, Shu-Hao Hu, Xiao-Bo Guo et al. · 0 citations
#machine learning Preprint Oct 2026

Minimax Optimal Regret for Causal Logistic Bandits with Counterfactual Fairness

We study causal logistic bandits with counterfactual fairness constraints. The causal structure is given through known factual and counterfactual feature maps that share an unknown logistic reward parameter, but the learner observes only factual rewards. Consequently, the directions determining counterfactual feasibili...

Junhyuk Huh, Seoungbin Bae, Dabeen Lee · 0 citations
#machine learning Preprint Sep 2026

From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences

In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intricate analysis. In th...

Yi-Bo Wang, Wen-Hao Yang, Si-Fan Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.