Skip to content

Category

machine learning

3,595 papers

Locality-Aware Redundancy Pruning for LLM Depth Compression

It is shown that inter-layer redundancy can be either localized or globally distributed depending on the LLM architecture, and Representation Locality Score (RLS) is introduced, derived from global inter-layer hidden-state similarity.

Vincent-Daniel Yun, Youngrae Kim, Woosang Lim et al. · 1 citation · ⚡1
#artificial intelligence Preprint May 2026

When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer Selection

Student-Centric Answer Sampling (SCAS) is proposed, a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost and is derived by a token-wise gradient decomposition and used to guide answer selection during training.

Zhengyu Hu, Zheyuan Xiao, Linxin Song et al. · 0 citations

Generalist Graph Anomaly Detection via Prototype-Based Distillation

ProMoS is introduced, the first unsupervised generalist GAD framework, which detects anomalies by modeling the abundant normality in unlabeled data, and proposes prototype-guided soft-label distillation to align teacher and student in a shared prototype space, enhancing cross-graph generalizability.

Yiming Xu, Zihan Chen, Z. Peng et al. · 0 citations
#machine learning Preprint May 2026

World Model Control by Trajectory Reachability Metrics

A fixed encoder is studied in a fixed encoder and trajectory reachability metrics (TRM), a small temporal pairwise cost trained from logged trajectories and used to rank predicted endpoints of a candidate action sequence against a goal is introduced.

Lian Li, Shengzhi Wang, Li-Bin Qiu et al. · 2 citations
#machine learning Preprint May 2026

Emulating the Forced Response of Climate Models with Generative Machine Learning

This research demonstrates that the model, ArchesClimate -- SSP, does not simply imitate scenarios seen during training, but is actually capable of modeling the response of a climate state to diverse forcings, an important step towards reliable and rapid climate model scenario generation.

Graham Clyne, Julia Kaltenborn, Peer Nowack et al. · 0 citations

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling

A dynamic per-layer scalar derived by adapting the LARS/LAMB trust-ratio principle to the orthogonalized setting, where the standard denominator candidates---the raw momentum norm or the polar-factor norm---either live in the wrong unit space or carry no update-scale information.

Yuxuan Lou, Yang You · 1 citation

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

AOPD replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning and maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.

Nan Jia, Haojin Yang, Xing-Chen Ma et al. · 18 citations · ⚡5
#artificial intelligence Preprint May 2026

Concepts Whisper: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations

The results suggest transformers rotate semantic content into spectrally quiet regions during contextualized processing, where, in some architectures, interventions may reduce grammatical disruption relative to high-variance steering.

Pratyush Acharya, Nuraj Rimal, H. Dhakal · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.