A depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons is proved, and the first exponential separation for ReLU networks between two fixed depths whose shallower network has depth at least $3 is proved.
CatchBench puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST), which none scores all three under one task-method interface.
Across various benchmarks, it is shown that the representations learned by Mol-JEPA deliver strong performance, demonstrating the value of incorporating biochemical context through latent space prediction.
Florian Rottach, Sebastian Schieferdecker, William Rudman et al.· 0 citations
This work proposes CoTeach, a Confidence-aware dual-teacher learning framework that dynamically selects the more reliable teacher for each node, and demonstrates that CoTeach consistently improves few-shot node classification performance while reducing unnecessary LLM utilization and associated monetary costs.
Hojin Kim, Sujin Yoon, Sungsu Lim et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally identifiable information (PII) the instruct model had already memorized. We first confirm that instruct models have already memorized PII but leave them latent, rarely surfacing one when asked. We then apply RL on benign factual data that contains no PII of any kind, and re-probe: a targeted probe over name->email pairs, and an untargeted free-recall prompt that simply asks the model to list the addresses it knows. PII extraction rises sharply under both: on DeepSeek-V3.1, verbatim recall@k increases from 0.155 to 0.370, a 2.4x gain. The effect scales with model size: across three models spanning 8B to 671B parameters, absolute leakage is largest in the biggest model. Meanwhile model's reasoning abilities and refusal rates are retained, indicating that RL selectively changes which memorized information is accessible rather than broadly altering the model. In summary, memorized private data can be made markedly more extractable by training that never touches it. This gives an adversary a route to memorized data that requires no privacy-relevant training signal and no access to the data itself -- only the ability to fine-tune on something innocuous.
In-cell learning is introduced, a paradigm for writing new knowledge only within these cells, so that re-quantizing the served weights reproduces the released integer codes and scales exactly.
Zifeng Liu, Yaxin Lu, Xuanhan Wu et al.· 0 citations
This work introduces the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection.
Zhongwei Yu, Yan Song, Xue Yan et al.· 0 citations
An AIM framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual, and RADAR, which combines relativistic adaptive geometry, decoupled residual correction, and second-order momentum filtering to improve the update direction and momentum estimation.
Zhi-Xin Ren, Yao Lyu, Congrong Li et al.· 0 citations
This paper proposes World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck and accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines.
Xi-Yuan Yang, S. Sarwar, Jingru Cheng et al.· 0 citations
This work evaluates SymBuild in three construction domains: computer-aided design (CAD) assembly, Mini-Programs, and exact-fill packing, and test additional framework instantiations in all four domains, demonstrating that SymBuild is an effective, analyzable method for anytime verified construction.
Weight-space composition supports coarse, input- and format-conditioned functional statements -- not a universal merging-performance predictor, and not one that training-format evaluations can see.
This work presents Task-to-Model Optimization (T2MO), a data-driven methodology for optimizing model selection in production coding workflows, and describes the methodology, optimization objective, evaluation protocol, and governance loop in a form suitable for production deployment and future empirical study.
Srinivasan Manoharan, Junhua Zhao, Fang Tu et al.· 0 citations
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026