Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Sep 2026

BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

BranchShine-CR is introduced, a 25M-parameter model for multilingual transcription into the International Phonetic Alphabet that supports compact IPA recognition capabilities under limited compute budget, for applications in low-resource on-device pronunciation assessment.

Nikhil Navas, S. Chevtchenko, Talisson Damiao et al. · 0 citations
#machine learning Preprint Sep 2026

Generative Atmospheric Super-Resolution from Heterogeneous In Situ Observations through Composable Interfaces

This work forms this reconstruction problem as generative atmospheric super-resolution and introduces composable observation interfaces for conditioning a single pretrained 13-variable atmospheric diffusion model without retraining the underlying model.

Yang Xu, Dibyajyoti Chakraborty, Hai-Wen Guan et al. · 0 citations
#machine learning Preprint Sep 2026

Growth-Inspired Graph Generation and Inverse Design of Mechanical Lattices via Dot Matrices Database Augmentation and GCNN

Natural load-bearing and transport networks are not assembled in a single step; they emerge through a temporally ordered process of growth, branching, reinforcement, and loop formation. Inspired by this developmental logic, this work introduces a morphogenetic graph-generation framework for mechanical lattices in which...

Wei-Yun Xu, Jia-Mu Liu · 0 citations
#machine learning Preprint Sep 2026

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

Predictive Action Chunk Learning first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation.

Yan-Gang Ren, Yu-Jie Yan, Zi-Rui Li et al. · 0 citations
#machine learning Preprint Sep 2026

Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization

This work regularizes the energy of each rank-one LoRA component, encouraging redundant components to vanish while preserving important ones, and proposes a principled rank-allocation method based on classical sparsity-inducing technique in signal processing and statistics.

Ze-Bang Xie, Chuan-Yang Zheng, Yik-Chung Wu et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Spectral Graph Neural Networks with Hermite Polynomials: A Comprehensive Study

We study spectral graph neural networks built from Hermite polynomials and propose HermNet, a simple model that combines a nodewise predictor with normalized Hermite propagation. Its sparse recurrence requires neither eigendecomposition nor a learned basis. We distinguish the basic model from optional coordinate calibr...

Shuang Wu · 0 citations
#machine learning Preprint Sep 2026

Response-state Learning for Transferable Vibrational Spectroscopic Characterization with Electron Prior

Vibrational spectral prediction can become inaccurate when localized stereoelectronic environments perturb intermediate response states and high-risk response units dominate characteristic spectral fingerprints, making prediction across external chemical space difficult. SO(3) Equivariant Neural Kalman Networks (SENK)...

Ze-Tong Li, Zhuo-Song Xie, Heng-Yu Fan et al. · 0 citations
#machine learning Preprint Sep 2026

Image Fidelity is Not Field Fidelity: Joint Thermodynamic Reconstruction and Error Localization in Neural Tomography

Neural fields for scientific tomography are optimized from 2D images, but the actual quantity of interest is often a latent 3D physical field. Because the forward map is many-to-one, low 2D image error need not certify a correct 3D field. Moreover, the latent field is not directly supervised during training, and its er...

Alan Hsu, Jenna Samra, A. Paraschiv et al. · 0 citations
#machine learning Preprint Sep 2026

LastOPD: Taming Collapse in Latent On-Policy Distillation

LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.

Jie Yang, Zheng-Yu Fang, Ze-Lin Xu et al. · 2 citations · ⚡1
#machine learning Preprint Open access Sep 2026

When Does Unsupervised Learning Succeed or Fail? A PoS Perspective on Reconstruction-Based Anomaly Detection

Reconstruction-based unsupervised learning can fail in two opposing ways: a model may reconstruct anomalies too accurately or discard valid nominal variation. Using the Pursuit of Subspaces hypothesis, we characterize these failures through the meet, union, and join geometries induced by the nominal components. Excess...

Mehmet Yama\c{c}, Yagmur Mustu, Muhammad Numan Yousaf et al. · 0 citations
#machine learning Preprint Sep 2026

Stream Recursion Model (SRM)

Mechanistic interpretability seeks to make verifiable statements about the internal behavior of large language models (LLMs). Many interpretability techniques struggle to scale with the increasing size and depth of architectures. Our solution to this is to introduce smaller models with structures that lend themselves t...

Asael Sorensen, Charles Brock, D. Chamberlain et al. · 0 citations
#machine learning Preprint Sep 2026

Monitoring Urban Traffic Dynamics at Fine Spatiotemporal Resolution Using Distributed Acoustic Sensing and Deep Learning

This study develops a deep learning-empowered analytical framework that converts raw ground vibration waveforms into spatiotemporal representations, detects vehicle trajectory, and infers traffic states from aggregated traffic volume and speed.

Hao Tian, Heng Cai, Xiao-Wei Chen et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.