Aug 2026· Mathematics· Vol 14, pp. 3130· 0 citations· 18 references
TL;DR
On two real Bayesian posteriors scored against long NUTS references, uniform pooling improved on the sample-count-matched single run under marginal JS, energy distance, and MMD2 in every recorded run, while the aggregation’s deficit is confined to the biased marginal metric.
Abstract
Normalizing-flow Markov chain Monte Carlo (MCMC), such as flowMC, augments local moves with a learned global flow proposal; a natural reliability idea is to pool samples from several such samplers. On 6 targets in 20 and 50 dimensions, a 5-member diverse flowMC ensemble appeared to reduce the average marginal Jensen–Shannon (JS) distance by about 18% relative to a single flowMC run. Matching the returned sample counts—our “equal budget”: wall-clock costs differ and are reported separately—reverses this verdict: the ensemble draws five times as many samples, and even a perfect sampler’s histogram JS estimate has a closed root-mean-square finite-sample floor c(1/N+1/M), with c a fitted coefficient set by the binning, which explains the apparent gain almost entirely. At a matched 10,000-sample budget, a single flowMC run matches or beats this ensemble on all eleven configurations and runs about six times faster. Applied to exact draws, the uncorrected reweight-and-resample aggregation reproduces about 93% of the ensemble’s elevation above the floor, separating a genuine law shift (coverage and tempering reweighting) from an estimator-level loss (resampling). On two real Bayesian posteriors scored against long NUTS references, uniform pooling improved on the sample-count-matched single run under marginal JS, energy distance, and MMD2 in every recorded run (a descriptive comparison at three and two repetitions), while the aggregation’s deficit is confined to the biased marginal metric. We recommend matched sample counts and floor reporting for histogram divergences, unbiased joint metrics alongside, and uniform pooling instead of uncorrected reweighting and resampling.
Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted pr...
A. Kurennoy, R. Yarullin, Fergal Reid· 0 citations
Benford's law specifies a marginal distribution for significant digits, whereas the usual first-digit Pearson $p$-value is calibrated under an independent multinomial sampling model. We separate these statements with four constructions that share the same one-time or pooled Benford target but have different joint struc...
Real-generation experiments that account for the pilot synchronization barrier show that HBN-async mitigates its overhead, helping translate statistical efficiency into practical evaluation benefits, and outperforming hindsight-tuned empirical and independent Bayesian baselines.
Qian-Li Shen, Xiang Li, Ruo-Meng Ding et al.· 0 citations
Batched multi-armed bandits update on a service's own schedule, and the usual implementation carries each arm's absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thompson Sampling (OR-TS...
Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is...
We study approximate sampling: given $N$ independent samples from a proposal distribution $\mu$, the goal is to select one whose distribution is close to a target $\pi$ specified only up to a normalizing constant. Block and Polyanskiy (2023) provide finite budget error bounds for approximate rejection sampling (RS) as...