Skip to content

Category

machine learning

5,991 papers

#artificial intelligence Book Feb 2024

SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning Systems

This study unveils the capability of attackers to generate adversarial policies even when restricted to partial observations of the victims in multi-agent competitive environments, and proposes a novel black-box attack (SUB-PLAY) that incorporates the concept of constructing multiple subgames to mitigate the impact of partial observability.

Oubo Ma, Yuwen Pu, L. Du et al. · 16 citations

Amortizing intractable inference in large language models

This work interprets chain-of-thought reasoning as a latent variable modeling problem and demonstrates that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization.

Edward J. Hu, Moksh Jain, Eric Elmoznino et al. · 110 citations · ⚡19
#machine learning Open access Oct 2023

Expected flow networks in stochastic environments and two-player zero-sum games

This work shows that EFlowNets outperform other GFlowNet formulations in stochastic tasks such as protein design and extends the concept of EflowNets to adversarial environments, proposing adversarial flow networks (A FlowNets) for two-player zero-sum games.

Marco Jiralerspong, Bilun Sun, Danilo Vucetic et al. · 11 citations · ⚡1
#machine learning Open access Oct 2023

Delta-AI: Local objectives for amortized inference in sparse graphical models

A new algorithm for amortized inference in sparse probabilistic graphical models (PGMs) is presented that enables off-policy training but avoids the need to instantiate all the random variables for each parameter update, thus speeding up training considerably.

J. Falet, Haebeom Lee, Esmeralda S. Whitammer et al. · 9 citations
#machine learning Open access Oct 2022

GFlowNets and variational inference

This paper builds bridges between two families of probabilistic algorithms: (hierarchical) variational inference (VI), which is typically used to model distributions over continuous spaces, and generative flow networks (GFlowNets), which have been used for distributions over discrete structures such as graphs. We demonstrate that, in certain cases, VI algorithms are equivalent to special cases of GFlowNets in the sense of equality of expected gradients of their learning objectives. We then point out the differences between the two families and show how these differences emerge experimentally. Notably, GFlowNets, which borrow ideas from reinforcement learning, are more amenable than VI to off-policy training without the cost of high gradient variance induced by importance sampling. We argue that this property of GFlowNets can provide advantages for capturing diversity in multimodal target distributions.

Esmeralda S. Whitammer, S. Lahlou, T. Deleu et al. · 120 citations · ⚡9
#machine learning Preprint Aug 2026

Implementing neural network mixed-effects models in Template Model Builder (TMB)

A general framework for implementing NMMs using Template Model Builder (TMB), which automatically integrates out random effects and evaluates the marginal objective function alongside its exact gradients, eliminating the need for manual derivations or ad hoc approximations.

Nan Zheng, H. Cheung, Vibhu Sharma et al. · 0 citations
#machine learning Preprint Aug 2026

"Train classical, deploy quantum"requires rethinking generalization

The results indicate that a converged moment-matching loss is not a reliable measure of generalization, and that train-classical, deploy-quantum workflows will need approaches that target generalization directly, leaving open whether better training objectives suffice or whether the model architectures themselves must change.

S. Raj, Natansh Mathur, A. Perdomo-Ortiz · 0 citations
#artificial intelligence Preprint Aug 2026

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

From an industrial code-generation improvement effort, a maintainer's perspective on why this work is hard in practice is offered, distilling three recurring challenges, zero-sum mixture design, yield as the binding metric, and end-to-end integration under uncertainty, and arguing that progress depends less on one-off recipes than on an engineering discipline for programming dataware.

Gopi Krishnan Rajbahadur, A. M. Ebrahimi, Boyuan Chen et al. · 0 citations
#machine learning Preprint Aug 2026

Minimax bounds for watermarked and masked recursive discrete distribution estimation

This work provides a lower bound that shows that it is impossible to improve performance by adding watermarks unless the false negative rate of detection also vanishes, and shows that in most regimes, the worst-case losses of a sequence of simple deterministic estimators match the corresponding lower bounds up to constants.

Millen Kanabar, Michael Gastpar · 0 citations
#machine learning Preprint Aug 2026

Segmentation of Bovid Dentition Under Imperfect Annotations: A Comparative Study of Convolutional and Attention Models

This paper presents a comparative study of segmentation architectures, ranging from convolutional backbones to vision transformers, applied to the B.O.V.I.D. dataset, a corpus of high-resolution bovid dental photographs paired with hand-made segmentation masks not originally designed for ML-based training, and evaluates a range of preprocessing and alignment techniques to mitigate the resulting label imperfections.

Keith G. Mills, Evan B. Sanders, Gregory J. Matthews et al. · 0 citations

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.