Skip to content
Preprint

On the Estimation of Chernoff Information

Aug 2026 · 0 citations · 22 references
Computer Science Mathematics

TL;DR

This work proves the $L_2$-consistency of the derivative estimator under mild regularity conditions on the densities and their domain and reformulates this optimization via a derivative condition, whose zero locates the optimal mixture parameter, and estimates the derivative directly using a-nearest-neighbor method.

Abstract

Chernoff information is a fundamental divergence measure characterizing the optimal error exponent in Bayesian binary hypothesis testing, with applications in information fusion, time-series analysis, and statistical learning theory. However, closed-form expressions exist only for simple parametric families, and nonparametric estimation remains difficult because the quantity is defined as an optimization of the unnormalized R\'enyi divergence over its order. We reformulate this optimization via a derivative condition, whose zero locates the optimal mixture parameter, and estimate the derivative directly using a $k$-nearest-neighbor method. We prove the $L_2$-consistency of the derivative estimator under mild regularity conditions on the densities and their domain. Coupled with a bisection procedure that locates the optimal parameter up to arbitrary precision, this yields an estimator for Chernoff information.

View source

Similar papers

NONPARAMETRIC BAYESIAN ESTIMATION

Frank Van Der SHOTA GUGUSHVILI, M. Meulen, Schauer Peter et al. · 1 citation
Preprint Aug 2026

On the minimax-rate optimality of approximate Bayesian computation in nonparametric problems

Approximate Bayesian computation (ABC) replaces likelihood evaluation in conventional Bayesian computation by simulation and comparison of observed and synthetic data. We demonstrate that ABC can be minimax-rate optimal in nonparametric settings. Our main result is a general contraction theorem for ABC posteriors based on summary statistics sieves and a localized prior-mass condition. We apply this theorem to Gaussian sequence estimation over Sobolev ellipsoids and to density estimation over bounded Sobolev-type classes and, under model-specific conditions, construct ABC procedures whose ideal posteriors and posterior means attain the corresponding minimax rates. Conditional on sampling from the specified priors, we show that the conventional Monte Carlo rejection ABC algorithm inherits the same rates when the number of simulation proposals grows at a sufficiently large exponential rate in the effective dimension.

H. Nguyen · 0 citations
Preprint Aug 2026

A Structural Characterization of Entropy Functionals

Entropy functionals and their associated divergences underlie many statistical methods, including maximum entropy inference, minimum divergence estimation, and goodness-of-fit testing, yet choosing among Shannon, R\'enyi, Tsallis, and more general entropies is often a matter of convention rather than structural principle. We introduce a measure theoretic framework in which admissibility requires the entropy of an input measure to be bounded above by that of its reference measure whenever the former is absolutely continuous with respect to the latter. Under generalized mean-value composition, we characterize all such entropy functionals and obtain a four-level hierarchy determined successively by the mean generator, entropy scale, and additivity assumptions. A continuous strictly monotone generator $g$ is admissible exactly when $t\mapsto g(1/t)$ is strictly convex for increasing $g$, or strictly concave for decreasing $g$. This resolves a question posed by R\'enyi (Proc. 4th Berkeley Sympos. Math. Statist. Prob., 1961) concerning which generalized means may replace the arithmetic mean in his entropy axiomatization. The same criterion is equivalent to strict convexity of an associated Csisz\'ar $f$-divergence generator and therefore yields data processing under Markov kernels with an exact equality condition. Within this hierarchy, product additivity singles out the R\'enyi family, while internal additivity, or product additivity together with arithmetic mean-value composition, singles out Shannon entropy. The characterization is constructive and yields new admissible entropy and divergence families, including integral-transform examples.

D. Lazarev · 0 citations
Preprint Aug 2026

A Bayesian Proof of the Bernoulli Theorem

We give a new proof of the Bernoulli theorem, conjectured by Talagrand and proved in the seminal work of Bednorz and Lata{\l}a. Our approach is based on information-theoretic ideas: lower bounds on the supremum of a Bernoulli process are translated to the fundamental limits of Bayesian estimation in a Cauchy additive channel. This leads to a new information-theoretic functional that characterizes Bernoulli-process suprema and plays a role analogous to Fernique's majorizing-measure functional for Gaussian processes. The same viewpoint yields a distributional strengthening: for any prescribed law of the index, we characterize the largest expected value attainable over all couplings of that index with the Bernoulli process. This extends to Bernoulli processes a phenomenon previously understood for Gaussian processes through the work of Fernique and Talagrand.

Jingbo Liu, Ilias Zadik · 0 citations
Preprint Jul 2026

A New Look at the Classical Estimation Problem

Bahadur's \emph{Lectures on the Theory of Estimation} develop the classical theory of point estimation inside the geometry of Hilbert space, and they record with unusual honesty where the theory strains: the locally best unbiased estimate depends on the parameter, a two-point parameter space yields an estimate Bahadur calls absurd, the odds ratio in binomial sampling has no unbiased estimate, and the virtues of maximum likelihood enter as heuristics and remain heuristics. We present a subset of the lectures, in Bahadur's notation and development, and at each strain make one small modification: for each value in the sample space, an estimate $\tau$ becomes a function on the parameter space rather than a point in it, the continuum of null hypotheses that Fisher described in 1955. Bahadur's own definition of an estimate, square-integrable at every distribution in the family, already supplies the domain. The payoffs are tracked lecture by lecture: estimators that exist at boundary samples where point estimates do not; an elementary lemma showing that no pointwise criterion admits a uniformly optimal estimator, which explains why admissibility, minimaxity, Bayes averaging, and unbiasedness arose as responses; assessment by information, $\Lambda(\tau)$, with the score attaining the Fisher information bound uniformly by a three-line argument; Cram\'{e}r--Rao attainment and sufficiency recovered as equality cases of that bound under two maps from point estimators to generalized estimators; and the maximum likelihood heuristics converted into exact statements about the score. Nothing classical is overturned; the classical apparatus is explained using Fisher's characterization of estimation as a continuum of significance tests.

Paul W. Vos · 0 citations
Open access Aug 2026

Optimality Notions for Resolvent Monte Carlo

Resolvent Monte Carlo estimates eigenvalues of large matrices by sampling Markov chains and reading the target value off a truncated resolvent quotient, trading exact arithmetic for a stochastic error that the almost-optimal sampling scheme is designed to suppress. This paper studies when that error vanishes outright. An exact closed-form identity is derived for the variance of the moment estimators of a general, possibly signed matrix, and is used to isolate a hierarchy of zero-variance notions ranging from the most local, which constrains only the first draws, through the finite-truncation regime that a practical run can certify, to the global regime in which every moment estimator is deterministic. Determinism of the estimator is separated from correctness of the eigenvalue it reports, and the exact conditions under which each notion holds are exhibited, together with the examples that separate them. A single edgewise condition, termed the eigen-triple condition, forces the truncated quotient to equal the target eigenvalue in finite samples; the associated moment and quotient variances are second order in the maximal edge defect and vanish at the eigen-triple. A linear-time procedure certifies the condition.

Tsvetelina Kostadinov, I. Dimov · 0 citations