Skip to content

Category

machine learning

3,595 papers

#machine learning Preprint Aug 2026

Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks

A depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons is proved, and the first exponential separation for ReLU networks between two fixed depths whose shallower network has depth at least $3 is proved.

Itay Safran · 0 citations
#machine learning Preprint Aug 2026

CatchBench: When Can an Agent Failure Be Caught?

CatchBench puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST), which none scores all three under one task-method interface.

Yue Zhao · 0 citations
#machine learning Preprint Aug 2026

Who Should Teach? Confidence-Aware Dual-Teacher Learning for Few-Shot Node Classification on Text-Attributed Graphs

This work proposes CoTeach, a Confidence-aware dual-teacher learning framework that dynamically selects the more reliable teacher for each node, and demonstrates that CoTeach consistently improves few-shot node classification performance while reducing unnecessary LLM utilization and associated monetary costs.

Hojin Kim, Sujin Yoon, Sungsu Lim et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data

Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally identifiable information (PII) the instruct model had already memorized. We first confirm that instruct models have already memorized PII but leave them latent, rarely surfacing one when asked. We then apply RL on benign factual data that contains no PII of any kind, and re-probe: a targeted probe over name->email pairs, and an untargeted free-recall prompt that simply asks the model to list the addresses it knows. PII extraction rises sharply under both: on DeepSeek-V3.1, verbatim recall@k increases from 0.155 to 0.370, a 2.4x gain. The effect scales with model size: across three models spanning 8B to 671B parameters, absolute leakage is largest in the biggest model. Meanwhile model's reasoning abilities and refusal rates are retained, indicating that RL selectively changes which memorized information is accessible rather than broadly altering the model. In summary, memorized private data can be made markedly more extractable by training that never touches it. This gives an adversary a route to memorized data that requires no privacy-relevant training signal and no access to the data itself -- only the ability to fine-tune on something innocuous.

Renfei Zhang, Niloofar Mireshghallah · 0 citations
#machine learning Preprint Aug 2026

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

This work introduces the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection.

Zhongwei Yu, Yan Song, Xue Yan et al. · 0 citations
#machine learning Preprint Aug 2026

Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization

An AIM framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual, and RADAR, which combines relativistic adaptive geometry, decoupled residual correction, and second-order momentum filtering to improve the update direction and momentum estimation.

Zhi-Xin Ren, Yao Lyu, Congrong Li et al. · 0 citations
#machine learning Preprint Aug 2026

Scaling Automatic Research Agents via World Models

This paper proposes World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck and accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines.

Xi-Yuan Yang, S. Sarwar, Jingru Cheng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified Construction

This work evaluates SymBuild in three construction domains: computer-aided design (CAD) assembly, Mini-Programs, and exact-fill packing, and test additional framework instantiations in all four domains, demonstrating that SymBuild is an effective, analyzable method for anytime verified construction.

Yi Liu · 0 citations
#machine learning Preprint Aug 2026

Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing

This work presents Task-to-Model Optimization (T2MO), a data-driven methodology for optimizing model selection in production coding workflows, and describes the methodology, optimization objective, evaluation protocol, and governance loop in a form suitable for production deployment and future empirical study.

Srinivasan Manoharan, Junhua Zhao, Fang Tu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.