Skip to content
Preprint

From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization

Sep 2026 · 0 citations
Mathematics

TL;DR

A theory is developed that connects a computationally tractable nonconvex procedure directly to statistical inference for nonlinear latent factor models with exponential-family links and partially observed entries with severe missingness, weak low-rank signals, and diminishing local curvature.

Abstract

Generalized latent factor models provide a flexible framework for analyzing high-dimensional non-Gaussian data, but principled estimation and uncertainty quantification under missingness remain substantially less developed. We develop a theory that connects a computationally tractable nonconvex procedure directly to statistical inference for nonlinear latent factor models with exponential-family links and partially observed entries. Our procedure combines a link-aware double-SVD initialization, a unilateral refinement that achieves rowwise consistency, and vanilla gradient descent. We show that the refined initializer enters a region of incoherence and contraction and that gradient descent remains in this region through implicit regularization, contracting rapidly down to the statistical estimation error without explicit incoherence or balancing regularization. Our central result is a uniform rowwise linear approximation for the actual output of gradient descent that isolates the leading score fluctuations from higher-order estimation and optimization errors. These expansions yield asymptotically valid individual and Gaussian multiplier-bootstrap simultaneous inference for latent factors, together with simultaneous confidence bands for missing-entry means, without requiring an additional debiasing step. The resulting estimation rate matches a restricted-class minimax lower bound up to logarithmic factors, while the theory accommodates severe missingness, weak low-rank signals, and diminishing local curvature. Simulations support the theoretical findings, and an application to large language model evaluation illustrates uncertainty-aware estimation and ranking of latent model capabilities.

View source

Similar papers

Preprint Sep 2026

Marginal maximum likelihood estimation and asymptotic theory for latent variable models in high dimensions

This work addresses a longstanding gap in the statistical foundations of marginal maximum likelihood estimation for high-dimensional latent variable models. Marginal maximum likelihood estimation is widely used to fit latent variable models across the social sciences, ecology, and machine learning. Despite its broad us...

Cheng-Yu Cui, Gong-Jun Xu · 0 citations
Preprint Sep 2026

Estimation and Inference for Latent Markov Models by Fourier Recursions

This paper introduces a unified truncated implementation that ensures uniform error control and numerical stability, preventing approximation errors from accumulating through the recursion, and provides the first general asymptotic theory for feasible approximate maximum likelihood estimation in LMMs.

Yan-Qi Huang, Chen-Xu Li, Qi-Wei Yao · 0 citations
Preprint Aug 2026

Duality and Error for Predictively Oriented Inference

This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a uni...

Aurya Javeed, D. Kouri, Teresa Portone et al. · 0 citations
Preprint Sep 2026

Identification of Nonlinear and Dependent Latent Factor Structure through Clique Search

Learning the structure of latent factor models involves two central challenges: (1) estimating the number of latent factors and (2) learning the support of the mapping from latent variables to observed variables. This is especially challenging for nonparametric regimes and nonlinear settings. We propose a method for la...

D. Kim, Qing Zhou · 0 citations
Preprint Aug 2026

A Two Stage Quasi-Likelihood Estimation Method for High Dimensional Generalized Structural Equation Models

Estimating high dimensional Generalized Structural Equation Models presents severe computational challenges. Traditional simultaneous estimators frequently suffer from numerical instability and prohibitive computational costs. Moreover, there are no tractable algorithms for families such as Poisson, negative binomial,...

M. Hattab · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.