Skip to content

Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling

Sep 2026 · 0 citations
Computer Science Engineering

TL;DR

The paradigm from isolated losses to iterative optimization over canonical models: population minimizers of standard conditional matching for terminal densities and approximate updates and define critical-point consistency as vanishing tangential displacement if and only if $\rho=\pi$.

Abstract

We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. The target is $\pi\propto\mu e^{\tau r}$, where $r$ is the reward, $\tau>0$ the inverse temperature, and $\mu$ denotes the pretrained model's terminal density for fine-tuning or the constant $1$ for sampling. We shift the paradigm from isolated losses to iterative optimization over canonical models: population minimizers of standard conditional matching for terminal densities. Under compatible smooth-realization assumptions, canonical velocities form a manifold diffeomorphic to the density manifold. Transporting the Fisher-Rao metric and mixture connection to this manifold, we show that the reverse-KL Hessian equals the metric, so the Newton direction coincides with the negative Fisher-Rao gradient. At terminal density $\rho$, each stage takes a tangential step generated by the regularized reward $r-\frac1\tau\log(\rho/\mu)$, followed by terminal-density-preserving canonicalization. This canonical retraction yields an exact finite-stepsize density characterization. For the ideal iteration, we prove strict reverse-KL descent away from the target for $0<\eta \le \tau$, global convergence under mild conditions, and local quadratic convergence for full steps ($\eta=\tau$). Covariance and gradient forms, each with forward or reverse regression-pair constructions, yield sample-wise tangential-update losses with the same population minimizer, without importance sampling or full-trajectory backpropagation. We develop approximate updates and define critical-point consistency as vanishing tangential displacement if and only if $\rho=\pi$. We recover representative methods as exact realizations, critical-point-consistent approximations, or objective-altering variants, enabling modular algorithm design. Our work advances the theory and algorithms of reinforcement learning for generative models.

View source

Similar papers

Preprint Sep 2026

Oracle high-dimensional $M$-estimation using smooth reparameterization for sparsity

This paper establishes a unified non-linear regularization framework for high-dimensional $M$-estimation, encompassing both linear models and Cox's proportional hazards models. Rather than relying on traditional additive non-convex penalties, the proposed paradigm embeds sparsity directly into the transformation for th...

Y. Nishiyama · 0 citations
#artificial intelligence Preprint Sep 2026

The Gaussian Is Enough: Flow-Matching Priors Do Not Help When Fine-Tuning Large Behavior Models

Modern robot imitation learning increasingly relies on generative policies based on diffusion or flow-matching models, which generate actions by transforming samples from a prior distribution. A key question is whether the choice of prior matters. Replacing the standard Gaussian with a closer-to-target, non-Gaussian pr...

Chen Xu, Rishi Shah, H. Kress-Gazit et al. · 0 citations
Preprint Sep 2026

Approximating Measures on Function Spaces: Transport and Truncation

This work introduces the class $\mathcal{P}_\psi(\mu)$ of measures that differ from a reference measure only through a finite-dimensional map $\psi$ while preserving the reference conditionals on its fibers, and develops approximation theory for fitting within it.

R. Baptista, Bamdad Hosseini, Alexander Hsu · 0 citations
Preprint Aug 2026

Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

It is shown that a symmetry fixes what variables are: a network layer is a sum over interchangeable units, so relabeling the units leaves it unchanged; given smoothness and the condition that a unit's gradient vanish at the origin, symmetry then enforces a universal leading form for the expansion about the near-zero we...

Zi-Yin Liu, Yi-Zhou Xu, Tomaso A. Poggio et al. · 0 citations
#machine learning Preprint Aug 2026

Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow

This work considers one-step generators from an optimal transport view, investigating Wasserstein Gradient Flow (WGF) for modeling smooth and controlled distributional evolution in probability space, and proposes a novel reward-guided fine-tuning of a one-step generative model via WGF.

Hoseong Hwang, Woorim Han, Joungin Chun et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.