Skip to content

Category

machine learning

3,595 papers

From AGI to ASI

How AI itself might continue to develop in a post-AGI world along the continuum of machine intelligence is investigated, which can intuitively be understood as a system that is more intelligent and cognitively capable than large organisations of humans.

Tim Genewein, Matija Franklin, Alexander Lerchner et al. · 4 citations

Twelve quick tips for designing AI-driven HPC workflows

This article offers a framework for transitioning from rigid execution pipelines to adaptive, intelligent computational environments, broadly applicable across distributed environments, they are particularly tailored to the resource-intensive throughput demands of modern computational biology.

J. Alnasir · 0 citations
#artificial intelligence Review Jun 2026

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

A dataset of 164 expert-annotated progress chains from the MIT PRIMES--Art of Problem Solving CrowdMath program (2016-2025), a collaborative research initiative whose discussions have led to peer-reviewed publications, is introduced.

Sherin Muckatira, Jesse Geneson, Slava Gerovitch et al. · 0 citations

Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference

Retrospective Harness Optimization is introduced, a self-supervised method that optimizes the agent harness using only past trajectories and alters the agent's behavior patterns and sustains higher accuracy during long-horizon sessions.

Wenbo Pan, Shujie Liu, Chin-Yew Lin et al. · 8 citations · ⚡1
#machine learning Preprint Open access Sep 2026

Don't Read Everything: A Curvature-Conditioned Query for Linear Attention

Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tasks. Existing remedies act on the write side of memory through gating, delta updates, or kernel feature maps, but the read step is left unchanged: every past key contributes additively to the output, so useful targets are diluted by the bulk of stored vectors. We borrow one specific piece of softmax's geometry to construct a cheap read-time contraction of the query. A second-order Taylor expansion of the softmax log-partition at the isotropic-attention point gives a local quadratic model whose curvature coincides with the running key covariance, a quantity that can be maintained with the same recurrent/chunkwise mechanism as the linear-attention state. The associated linear operator contracts the query along the high-variance directions of memory before it reads the state. We call this mechanism Curvature-Conditioned Query (CCQ). CCQ modifies only the read step and is composable with any linear-attention backbone. Attached to GLA and Gated DeltaNet, it improves perplexity, zero-shot downstream accuracy, S-NIAH retrieval at and beyond the training context, length-extrapolation perplexity from 4K to 20K, and LongBench accuracy.

Dong Le, Thong Nguyen, Cong-Duy Nguyen et al. · 0 citations

DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models

DASH is introduced, which supervises the conditional and unconditional branches independently and an anchor term regularises the conditional prediction toward ground-truth noise, and the teacher's final learned per-timestep curriculum transfers into the student as a frozen prior.

A. Shafi, Kazi Saeed Alam, Sk. Imran Hossain et al. · 1 citation
#artificial intelligence Preprint May 2026

Self-Correction Can Amplify Hallucinations: Fact-Level Repair with Graph-Based Evidence Routing in Multimodal Generation

TIGER is presented, an inference-time framework that redesigns feedback for localized repair that reduces unsupported content while preserving task quality and a CrisisFACTS case study suggests that the same repair mechanism can improve grounding in multi-source settings.

Kaixiang Zhao, Tianrun Yu, Shawn Huang et al. · 0 citations

Exploring Autonomous Agentic Data Engineering for Model Specialization

This study formalizes Autonomous Agentic Data Engineering, a novel task designed to evaluate LLMs as autonomous data engineers that drive model specialization through end-to-end data curation, and charts a path toward agent-driven model specialization.

Yujie Luo, Xiangyuan Ru, Jingsheng Zheng et al. · 2 citations
#machine learning Preprint Open access Sep 2026

Privacy-Enhanced Zero-Order Federated Learning via xMK-CKKS over Wireless Channels

Homomorphic encryption (HE) enables privacy-preserving aggregation in federated learning (FL) by allowing the server to operate on encrypted data without decryption. Existing HE-over-the-air (OTA) methods mainly rely on single-key HE schemes and require channel estimation or pre-equalization to compensate for wireless fading. However, single-key HE remains vulnerable to honest-but-curious (HBC) clients holding the shared secret key, while multi-key HE provides stronger client-level security by assigning each device its own secret key. We propose a four-phase protocol that enables the aggregation of xMK-CKKS over a shared wireless channel without channel estimation. The protocol retransmits partial public keys and ciphertexts through the same channel realization, so that the dominant large-modulus encryption terms cancel algebraically during decryption. We integrate this protocol with zero-order FL over slowly varying LoS-dominant channels, where each device transmits a single encrypted scalar per round and the communication/encryption overhead is independent of the model dimension. We show that the residual noise induced by encryption and wireless aggregation preserves the standard convergence rate \(O(1/\sqrt{K})\) up to a negligible noise floor, where $K$ is the number of communication rounds. The protocol assumes a non-trusted server and is secure against HBC clients, preventing any client from recovering the local updates of other participants. Numerical results on MNIST and CIFAR-10 validate the theoretical analysis.

Anthony Ayli, Khalil Harris, Jihad Fahs et al. · 0 citations

Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts

This paper presents a method for aggressively pruning experts from modern mixture-of-experts LLMs while incurring negligible degradation in translation quality, and shows that translation requires only a fraction of the LLM, enabling substantial compression of the MoE blocks that contain over 90% of parameters.

Liu O. Martin, Lucas Bandarkar, Nanyun Peng · 2 citations · ⚡1

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.