Skip to content

Category

machine learning

9,394 papers

#machine learning Preprint Open access Sep 2026

Assessing Predictive Models for Fairness Based on Activity-Space Patterns

Assessing the spatial fairness of predictive models involves establishing whether they are statistically penalizing (favoring) individuals associated with certain geographical locations. Literature on this topic makes the fundamental assumption that each individual is assigned to a single geographical location (e.g., p...

Francesco Lettich, Mario A. Nascimento, Chiara Pugliese et al. · 0 citations
#machine learning Preprint Open access Sep 2026

DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum

We study differentially private optimization with matrix-orthogonalized momentum. DP-Muon uses conventional global per-example clipping and one Gaussian gradient release per step; matrix updates and auxiliary updates are post-processing. Our main contribution concerns the additional mean distortion created when fresh G...

Jihwan Kim, Chenglin Fan · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Monotone Neural Policy Iteration for High-Dimensional First-Order Hamilton--Jacobi--Bellman Equations

We analyze a neural semi-discrete method for high-dimensional first-order Hamilton-Jacobi-Bellman (HJB) equations with known or learned dynamics. Centered differences and an artificial viscosity $Nh=O(h)$ define a monotone operator evaluated through $2d+1$ shifted network queries; policy iteration solves the resulting...

Minseok Kim, Yeongjong Kim, Namkyeong Cho et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Generalization Guarantees on Data-Driven Tuning of Gradient Descent with Langevin Updates

We study learning to learn through the lens of hyperparameter tuning. We propose the Langevin Gradient Descent Algorithm (LGD), which approximates the mean of the posterior distribution defined by the loss function and regularizer of a regression task with convex objective. For classification tasks, the LGD algorithm e...

Saumya Goyal, Rohith Rongali, Ritabrata Ray et al. · 0 citations
#machine learning Preprint Open access Sep 2026

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving

Mixture-of-Experts (MoE) models have become mainstream for scaling language models to hundreds of billions of expert parameters. Despite sparse expert activation, existing inference engines keep all experts GPU-resident, crowding out the key-value cache in large-batch, long-output offline workloads. We present FluxMoE,...

Qingxiu Liu, Yongchao He, Runhan Jiang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

Token-level sparse attention mechanisms, exemplified by DeepSeek Sparse Attention (DSA), achieve fine-grained key selection by scoring every historical key for each query through a lightweight indexer, then computing attention only on the selected subset. While the downstream sparse attention itself scales favorably, t...

Yufei Xu, Fanxu Meng, Fan Jiang et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Longitudinal Risk Prediction in Mammography with Privileged History Distillation

Longitudinal mammography screening has become an important source of information for improving future breast cancer risk prediction. However, the performance of current longitudinal mammography models degrades when prior examinations are unavailable at inference, creating a structured privileged-information setting in...

Banafsheh Karimian, Soufiane Belharbi, Alexis Guichemerre et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification

Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely debated and underexplored. We introduce HorizonMath, a benchmark of 113 predominantly unsolved prob...

Erik Y. Wang, Sumeet R. Motwani, James V. Roggeveen et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Partial GFlowNet: Accelerating Convergence in Large State Spaces via Strategic Partitioning

Generative Flow Networks (GFlowNets) have shown promising potential to generate high-scoring candidates with probability proportional to their rewards. As existing GFlowNets freely explore in state space, they encounter significant convergence challenges when scaling to large state spaces. Addressing this issue, this p...

Xuan Yu, Xu Wang, Rui Zhu et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Evaluating Memory Structure in LLM Agents

Modern LLM-based agents and chat assistants rely on long-term memory frameworks to store reusable knowledge, recall user preferences, and augment reasoning. As researchers create more complex memory architectures, it becomes increasingly difficult to analyze their capabilities and guide future memory designs. Most long...

Alina Shutova, Alexandra Olenina, Ivan Vinogradov et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Prediction--Loss Alignment for Sampler--Robust Flow Matching Training

Recent work has popularized a practical recipe in diffusion and flow matching: predict the clean signal $x$, convert it to a velocity, and train through a velocity-space loss. The conversion contains a singular endpoint amplification and therefore appears prone to unstable optimization, yet recent systems obtain strong...

Jiadong Hong, Lei Liu, Xinyu Bian et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Smoothing the Score Function to Enhance Generalization in Diffusion Models

Diffusion models achieve remarkable generation quality, yet face a fundamental challenge known as memorization, where generated samples can replicate training samples exactly. We develop a theoretical framework to explain this phenomenon by showing that the empirical score function (the score function corresponding to...

Xinyu Zhou, Jiawei Zhang, Stephen J. Wright · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.