Skip to content
#software testing Open access

The historical development of sample size estimation: from Huygens and Bernoulli to the present.

Aug 2026 · Journal of the Royal Society of Medicine · pp. 1410768261479691 · 0 citations · 38 references
Medicine

Abstract

This paper traces the historical development of frequentist sample size estimation from its philosophical origins to its present-day complexity. Preliminary concepts were identified by Christiaan Huygens' work on expected value and Jacob Bernoulli's law of large numbers, which first linked sample size and estimation accuracy. The 18th and 19th centuries brought major advances in probability theory through the work of Pierre-Simon de Laplace, Carl Friedrich Gauss and Siméon-Denis Poisson, yet explicit sample size planning remained uncommon. The early 20th century saw the emergence of methods for sample size calculations based on Jerzy Neyman and Egon Pearson's hypothesis testing framework and Sir Ronald Aylmer Fisher's experimental design principles. While Donald Mainland and Austin Bradford Hill referred indirectly to these as early as the 1930s, it took many decades before their explicit use became common. After the Second World War, contributions from figures such as Abraham Wald further embedded sample size planning with sequential methodologies. From the 1970s onward, standardised formulas, regulatory requirements, reporting standards such as CONSORT and statistical software consolidated frequentist sample size estimation as a routine component in applied research. In the 21st century, simulation-based, adaptive and Bayesian approaches, together with open-source computational ecosystems, have expanded the scope and accessibility of sample size methods. In contemporary research, sample size estimation has evolved into a multifaceted discipline; its methodological sophistication is contingent upon the underlying objective, illustrating the persistent divergence between explanatory inference and decision-oriented design.

Read PDF

Similar papers

Preprint Jul 2026

A New Look at the Classical Estimation Problem

Bahadur's \emph{Lectures on the Theory of Estimation} develop the classical theory of point estimation inside the geometry of Hilbert space, and they record with unusual honesty where the theory strains: the locally best unbiased estimate depends on the parameter, a two-point parameter space yields an estimate Bahadur calls absurd, the odds ratio in binomial sampling has no unbiased estimate, and the virtues of maximum likelihood enter as heuristics and remain heuristics. We present a subset of the lectures, in Bahadur's notation and development, and at each strain make one small modification: for each value in the sample space, an estimate $\tau$ becomes a function on the parameter space rather than a point in it, the continuum of null hypotheses that Fisher described in 1955. Bahadur's own definition of an estimate, square-integrable at every distribution in the family, already supplies the domain. The payoffs are tracked lecture by lecture: estimators that exist at boundary samples where point estimates do not; an elementary lemma showing that no pointwise criterion admits a uniformly optimal estimator, which explains why admissibility, minimaxity, Bayes averaging, and unbiasedness arose as responses; assessment by information, $\Lambda(\tau)$, with the score attaining the Fisher information bound uniformly by a three-line argument; Cram\'{e}r--Rao attainment and sufficiency recovered as equality cases of that bound under two maps from point estimators to generalized estimators; and the maximum likelihood heuristics converted into exact statements about the score. Nothing classical is overturned; the classical apparatus is explained using Fisher's characterization of estimation as a continuum of significance tests.

Paul W. Vos · 0 citations
Preprint Jul 2026

Why Constants Matter in Distribution Testing: From Uniformity to Calibration

Distribution goodness-of-fit testing has developed a powerful rate-level theory: we often know how the required sample size scales with the alphabet size, the separation from the null, and the target error probability. Uniformity testing is the canonical example. One can distinguish the uniform distribution on $N$ categories from alternatives at total-variation distance at least $\epsilon$ with far fewer than $N$ samples, and the optimal scaling is now well understood. But rate-level theory leaves an important question unresolved: among several tests with the same sample-complexity order, which one actually gives the best risk or power? This is a constant-level question. It is especially relevant in modern applications where distribution testing is used not merely as an asymptotic abstraction, but as a practical design tool. This note argues that sharp constants in distribution testing play a role analogous to Fisher information in parametric estimation and Pinsker's constant in nonparametric estimation. First, they distinguish between tests that are all rate-optimal but not equally powerful. Second, they reveal the effective signal-to-noise ratio governing the testing problem. Third, they can guide tuning-parameter choices in downstream applications. We illustrate this perspective through large-alphabet uniformity testing and then explain why the same logic matters for choosing the number of bins in calibration testing.

A. Kipnis · 0 citations
Jul 2026

Finite samples, overlapping returns and the illusion of precision in risk management

This study quantifies the material distortions in empirical return distributions caused by finite samples and overlapping returns, raising caution about the reliability of raw historical simulation for risk management. We aim to provide actionable guidance on when historical data can be informative and when it may be misleading. Using controlled simulations of a Geometric Brownian Motion – a best-case scenario of stationarity and ergodicity – we isolate the distorting effects of sample size (1–100 years) and overlapping windows on distributional accuracy, employing Kolmogorov-Smirnov, Anderson-Darling and other distance metrics. Even a century of daily data leaves the mean estimated with substantial relative uncertainty and tail estimates highly variable. Overlapping returns improve nominal sample size and estimates of central tendency but introduce dependencies that materially distort tail risk, inflating Anderson-Darling statistics by a factor of 14 in our simulations. The trade-off is purpose-dependent: overlapping windows help describe the distribution body but hinder tail-risk inference. The analysis relies on ideal conditions rarely met in real markets. This limitation is intentional: documenting distortions under these favorable conditions establishes a conservative lower bound. In practice, with non-stationary, path-dependent data, distortions are likely larger. Risk measures calibrated from raw historical return distributions should generally avoid overlapping data for tail-risk applications. Overlapping windows may be acceptable for descriptive visualization or central-shape estimation, provided that effective sample size is reported. Even the finite-sample results alone call for a move from naive historical extrapolation to prudence, combining historical estimates with scenario analysis, robustness checks and forward-looking distributions. This controlled lower-bound study shows that even under ideal conditions, finite samples and overlapping windows materially affect distributional estimation, especially in the tails, with clear implications for risk measurement. To our knowledge, it is the first to systematically quantify these distributional biases.

R. M. Gaspar · 0 citations
Preprint Aug 2026

Let Time Tell: Identification and Gaussian Process Estimation for Interrupted Time Series

We study causal inference in interrupted time series designs where a treatment affects every unit simultaneously, so that the contemporaneous controls used by difference-in-differences and synthetic control are unavailable and the counterfactual must be extrapolated from a unit's own pre-treatment history. We establish identification within the potential outcomes framework and estimate the counterfactual by Gaussian process regression. Rather than committing to a single best-fitting trend, the estimator retains the functions consistent with the pre-treatment series and widens its intervals where extrapolation magnifies their divergence. Connecting it to reproducing kernel Hilbert space theory, we derive a bias decomposition that isolates the component extrapolation inflates and a worst-case bound on that component, justifying the Gaussian process estimator's posterior variance as extrapolation-aware uncertainty quantification. In closed form, the band equals the worst-case divergence the model class permits among functions consistent with the pre-treatment data. The method is illustrated with calibrated simulations and an analysis of handgun purchases after the Supreme Court's Heller decision, a universal treatment whose practical effect concentrates in a single jurisdiction. An R package, gpss, implements the approach.

S. Cho · 0 citations
Preprint Jul 2026

Local permutation tests for conditional independence: an adaptive binning perspective

In this work, we study the problem of testing conditional independence between random variables $X$ and $Y$ given a confounder $Z$. The local permutation test (LPT) offers a principled approach to this problem by partitioning the $Z$-space into pre-specified bins, and permuting the $X$ and $Y$ data within each bin, to assess the significance of an observed test statistic. However, when the partitions are pre-fixed, the resulting partition can be poorly balanced, as some bins may contain most of the samples while others contain only a few. This motivates the use of data-adaptive binning strategies, such as equisized bins with a fixed (typically small) number of points. We study this natural and practically important extension of LPT, providing finite-sample bounds on the Type I error for an arbitrary test statistic, providing stronger validity results than previously known. We also show that LPT attains power comparable to the oracle likelihood ratio tests derived from the Neyman-Pearson lemma. Within a linear confounder model class, we further analyze the effect of bin size and demonstrate that constant bin sizes can match the performance of partitions with growing bin-size. These results, further supported by extensive numerical simulations, position the proposed data-adaptive strategy as both practically implementable and statistically efficient.

David Chen, Rohan Hore, R. Barber · 0 citations
Review Open access 2026

BUILDING A CLUSTER LINEAR REGRESSION WITH CONSTRAINTS ON THE SIZE OF SUBSAMPLES OF DATA

The paper provides a brief overview of publications related to the division of the initial data sample into disjoint fragments when modeling complex objects. In particular, there is considered distributed asymmetric least squares estimation based on a Poisson subsample; a new methodology for the structural synthesis of specialized parallel computational sub-blocks for implementing a group method of data processing algorithms; ANOVA class method based on a dedicated subsample of data, which suggests ways to reduce the effect of biased variance estimation; a model of the intensity of occurrence and size distribution of forest fires, combining the theory of extreme values and point processes within the framework of a new Bayesian hierarchical model; recent advances in theory, methods and implementations of quantile regression in the context of massive and streaming data; methods for reducing the size of the initial sample due to missing values. The authors formulate the task of identifying the parameters of a cluster regression model with predefined capacities of each allocated subsample of the initial data sample. When assigning the sum of the error modules of the approximation as a loss function for calculating the model parameters, this task is reduced to a linear Boolean programming problem. The case is considered when the sizes of the formed clusters are approximately equal. Two new variants of the cluster regression model for the development of the chemical industry in the Russian Federation have been constructed.

S. I. Noskov, V. A. Telenkov · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.