Skip to content
Preprint

Exact Rejection Sampling for Non-Gaussian State Space Models

Aug 2026 · 0 citations
Economics

TL;DR

This work constructs a proposal that dominates the target by a known constant, generally unavailable for non-Gaussian state space models, yielding independent exact smoothing draws and an unbiased likelihood estimator whose relative variance is at most $1/p-1$ per draw at acceptance probability $p$.

Abstract

Rejection sampling requires a proposal that dominates the target by a known constant, generally unavailable for non-Gaussian state space models. We construct such a proposal for the latent state path, yielding independent exact smoothing draws and an unbiased likelihood estimator whose relative variance is at most $1/p-1$ per draw at acceptance probability $p$. The method covers scalar states with affine Gaussian dynamics and log-concave observation densities, including multivariate observations. Transition twisting makes the log target-to-proposal ratio separable, and tangent-line twists make each term nonpositive, producing an attained, sharp dominating constant. With a companding node placement, the accumulated envelope error is $O(T/G^2)$ for a sample of length $T$ with $G$ nodes per date, so $G\propto\sqrt{T}$ keeps acceptance bounded away from zero; for stochastic volatility, the required conditions hold almost surely. A simpler mode-centered grid shows the same scaling empirically. At $T=2{,}000$, acceptance is $75\%$, versus roughly $10^{-16}$ for the Gaussian envelope.

View source

Similar papers

Preprint Jul 2026

LatentFlow: A General Framework for Conditioning Stochastic Processes

Stochastic-process models are, as a rule, far easier to simulate than to condition. Non-linear observations, non-Gaussian likelihoods, black-box information, and global constraints all induce intractable conditional laws, requiring bespoke, model-specific constructions. We introduce LatentFlow, a single framework for conditioning stochastic processes, with no learned neural approximations and no training. Our starting point is to write the stochastic process as the deterministic image of a tractable latent innovation, $f_0 = T_{\vartheta}(\xi_0)$, with $\xi_0$ sampled from a simple reference distribution. This reduces process-level conditioning to latent-space inference: pull the likelihood back through $T_{\vartheta}$, sample the resulting latent law with a tractable guided probability flow, and push the samples forward. This construction is provably exact at the level of the target law; in practice, approximation enters only through finite terminal noising, Monte Carlo guidance, and time discretisation of the continuous-time dynamics, each of which is explicit and systematically reducible. As LatentFlow is training-free, conditioning reduces to solving a single reverse-time SDE. This enables conditional sampling in seconds on a single desktop CPU across model classes that have never shared a scalable method: classical spatial priors, nonlinear stochastic dynamics, mechanistic models from the physical and life sciences, stochastic PDEs, heavy-tails and extremes, point and discrete-state processes, and neural or simulator-defined processes.

Louis Sharrock, L. Astfalck, Henry Moss · 0 citations
Preprint Aug 2026

Estimation of distribution functions, their jumps and interval probabilities under measurement error

We consider the classical additive measurement-error model $X=Y+Z$, where the latent random variable $Y$ has unknown distribution $F_Y$ and the error $Z$ has a known distribution. We develop direct estimators for three functionals of $F_Y$: (i) $F_Y(x)$ at continuity points; (ii) interval probabilities $F_Y(y)-F_Y(x)$ when $x<y$ are continuity points; and (iii) the size of a jump at a prespecified discontinuity. We derive non-asymptotic bias and variance bounds, and establish asymptotic unbiasedness and consistency. Unlike previous work, we do not require $F_Y$ to admit a density, have a mixture representation, or satisfy global Sobolev smoothness assumptions. The framework accommodates arbitrary latent distributions, including those with both discrete and continuous components, and distributions with multiple jumps. These results rely on a link between Fourier inversion theorems and the algebraic structure of a class of estimators proposed in Mynbaev, Martins-Filho and Henderson (2022). A simulation study evaluates feasible tuning procedures and, where available, compares the finite-sample performance of the proposed estimators with existing methods.

K. Mynbaev, Carlos Martins-Filho, Chad Brown · 0 citations
Preprint Aug 2026

Finite-Probe Total-Variation Certificates for Finite-Basis Drifting Models

Drifting objectives compare a target and model distribution through a vector field observed noisily at finitely many locations. We ask what distributional conclusion such a frozen measurement system warrants. For integrable antisymmetric interactions and absolutely continuous laws in a declared finite density basis, the unnormalized sampled numerator satisfies $\operatorname{vec}(V_X)=Mc$, where $c$ is an antisymmetric mismatch and $M$ is probe-dependent. This identity yields an a posteriori total-variation (TV) upper confidence bound accounting for held-out field noise, estimated-operator error, and externally validated $L^1$ residual radii around normalized density approximants in the span; a nonpositive observability margin returns the trivial TV bound and abstains. The audit recomputes this numerator from held-out samples; a normalized drift statistic requires a separate joint numerator--denominator analysis. For Gaussian-RBF interactions, a global envelope supports distribution-free and empirical-Bernstein radii without truncation, with companion bounds for the Laplace similarity in the original drifting objective. We characterize random-probe observability by a population Gram matrix, identify rank and symmetry degeneracies, and prove large-bandwidth collapse toward mean matching. Synthetic studies exercise Gaussian and Laplace numerators, separately prespecified bounded-vector and variance-adaptive radii, Monte Carlo-calibrated operators, nonzero residual radii around normalized finite-basis approximants, outward-rounded observability bounds, and designed abstention. A joint basis-size/dimension stress path extends evaluation through $m=8$. The result is a conditional diagnostic for a finite density class, or for normalized finite-basis density approximants with external residual radii, not a universal guarantee from small training drift.

Sam Andersson, Ricky Mol'en · 0 citations
Conference Jul 2026

Randomized Kaczmarz EM for Large-Scale State-Space Models

The Expectation–Maximization (EM) algorithm is a standard tool for estimating linear state-space models, but its reliance on dense covariance matrices makes it difficult to use in large dimensions. We show that, for systems with sparse or banded dynamics and simple noise structures, EM can be run without ever forming these matrices. The key step is to replace the usual smoother with randomized Kaczmarz iterations in information form, reducing the covariance memory cost to ${{\mathcal{O}}}(n)$ and yielding a total footprint of ${{\mathcal{O}}}(Nn)$ for the state trajectory, versus ${{\mathcal{O}}}\left({N{n^2}}\right)$ for the classical smoother. The overall behavior of EM is preserved: the likelihood still improves up to a tolerance set by the iterative solver, and shrinking this tolerance recovers the classical updates. Experiments on large advection–diffusion models illustrate the payoff: RK-EM matches standard EM at moderate sizes and continues to operate at n = 65,536, where traditional methods run out of memory. When the system is sparse and reasonably well observed, this matrix-free version offers a practical alternative to ensemble smoothers, with no localization or inflation parameters to tune.

J. Almutawa · 0 citations
Preprint Jul 2026

Mixing-Law Uncertainty in Multivariate Normal Mean-Variance Mixtures: Semi-parametric Estimation and Robust Cumulative-Prospect Decisions

The distribution of a normal mean-variance mixture depends on the law of its positive mixing variable. We compare six parametric mixing laws with a grid nonparametric maximum likelihood estimator under the same determinant identification constraint. The mixing mean $m=\E(Z)$ is estimated and is not fixed at one. A paired block bootstrap is used to compare multivariate holdout log scores. The models that cannot be distinguished from the model with the largest score define a finite ambiguity set. We then consider a cumulative prospect problem on a common portfolio direction. For each model in the set, the NMVM representation gives a scalar projected return and a corresponding prospect-value function of the exposure. The distributionally robust decision maximizes the lower envelope of these functions. We prove existence of a solution, give the candidate points for the piecewise smooth problem, derive a reference-gap scaling result, and construct an interval branch-and-bound certificate for the finite-scenario optimum. In an application to 30 stock returns, the mixture models give higher holdout density scores than the multivariate Gaussian model. Several parametric and semi-parametric models, however, remain in the ambiguity set. The worst-case model is therefore determined at the portfolio optimization stage rather than selected in advance from a point estimate of the holdout score.

Nuerxiati Abudurexiti · 0 citations
Preprint Jul 2026

The Boltzmann structure of sampling: Intrinsic $p$-value and its emergent closed-form expression

We consider observables $X$ whose realizations in a sampled dataset are restricted, for example by measurement resolution, to a finite set of distinguishable categories within their possibly infinite theoretical domain. Given the probabilities of observable categories, we study the coarse-grained probability mass of families of possible datasets generated through a forward sampling process. The combinatorial construction induces an intrinsically discrete $p$-value defined directly from the sampling process rather than through additional probabilistic structure on observables. Specifying the sample means of $d+k$ arbitrary functions $g_\alpha(X)$ defines a linear family of datasets whose probability mass is obtained as a weighted sum over integer lattice points contained within the associated polyhedron. To overcome the intractable large-$N$ combinatorics, we derive via saddle-point techniques a density approximating these probability masses in the continuum limit of forward sampling within the multinomial universality class. As a demonstration, we consider conditional sampling, where $d$ structural means are fixed while $k$ means vary over admissible datasets. The information geometry emerging from the saddle-point density, together with the spherical symmetry arising at large $N$ from the intrinsic $p$-value construction, enables efficient computation of the $p$-value in the Laplace approximation via the $\chi^2_k$ distribution. The resulting statistic is given by the semi-analytic expression $2N$ times the Kullback-Leibler divergence between the information projections associated with the corresponding structural and observed linear families. These projections can be computed efficiently via standard numerical routines converging for sufficiently well-behaved sample means.

Orestis Loukas · 0 citations