Skip to content

Category

machine learning

3,595 papers

Bergson: An Open Source Library for Data Attribution

Bergson is an open source library that aims to enable faster progress in the field by providing a host of techniques that scale to very large language models and pre-training datasets, and provides quality of life tools for researchers.

Lucia Quirke, Louis Jaburi, David O. Johnston et al. · 0 citations

When Design Rules Break: Benchmark Composition Determines Whether Label Informativeness Predicts GNN Aggregator Choice

The results suggest that benchmark composition, rather than numerical insufficiency, determines whether design rules appear to generalize, and that the Facebook-100 regime provides a concrete target for future adaptive aggregation methods.

Neha Sharma, Ritesh Sharma · 0 citations

Divide-and-Conquer Modeling for the CTF-4-Science Lorenz Benchmark

This submission documents the divide-and-conquer modeling strategy developed for the CTF-4-Science Lorenz Chaotic Systems Challenge at AI-DEEDS 2026, which shows that bounded, scenario-specific updates can outperform broad model replacement on mixed chaotic forecasting benchmarks.

Shun-Dong Li · 0 citations
#machine learning Preprint Jun 2026

QueryGraph: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation

This work introduces a system that converts natural language queries into structured graphs and executes them via a deterministic planner, which uses depth-first search to resolve dependencies and combine results across tools, improving reliability and enabling queries beyond traditional keyword-based search.

A. Chakravarthy, Vidhi Kulkarni, Duen Horng Chau · 0 citations

TriHead-GAN: A Generative Adversarial Network with Triple-Head Discriminator for Carbon Emission Time Series Generation

TriHead-GAN is proposed, a Transformer-based adversarial framework whose triple-head discriminator jointly supervises three complementary aspects of the joint distribution: distributional authenticity via a Wasserstein critic, cross-variable dependency via leakage-free regression of the target variable, and step-wise temporal smoothness via adjacent-difference prediction.

Ze-Sen Wang, Lijuan Lan, Yong-Gang Li et al. · 0 citations

GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution

This work reframe attribution as subset-level counterfactual utility prediction and introduces GRASP, an interaction-aware surrogate that more than doubles the task-level rank correlation for counterfactual subset fidelity while reducing upfront artifact construction costs by nearly an order of magnitude.

Yue Min, Rui Chen, Yujun Li · 0 citations

Mamba-Assisted Non-Markovian Closure for Reduced-Order Modeling

The Mamba-Assisted Closure (MAC) framework is proposed, which employs a Mamba-based sequence model to predict the closure from the resolved trajectory and couples the learned closure with the reduced-order governing equations through a numerical integrator to advance the resolved variables in time.

Zhifei Wei, S. Qadeer, Panos Stinis · 0 citations

GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning

GRZO is a Group-Relative Zeroth-Order optimizer that draws one pseudo-independent perturbation per mini-batch example and aggregates the per-example losses through group-relative normalization, raising the effective gradient-direction count from one to the batch size at no additional forward cost while preserving inference-level memory.

L. Tan, Yequan Zhao, Yifan Yang et al. · 0 citations

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

This work proposes a novel block-diagonal Riemannian metric derived from the pullback of the Frobenius inner product and develops a Riemannian gradient descent algorithm that uses a tuning-free Gaussian step size and scales linearly in the number of observed entries per iteration.

Pratik Jawanpuria, Ankish Chandresh, Bamdev Mishra · 0 citations

LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training

This work proposes LaRA, a layer-wise representation analysis framework for detecting contamination in RL post-trained LLMs and finds that contamination produces progressive geometric deviations across layers, including amplified perturbation sensitivity, stronger directional collapse, and enhanced local rigidity.

Minju Gwak, Minseok Kwak, Dongseok Lee et al. · 0 citations

OISD: On-Policy Internal Self-Distillation of Language Models

The OISD framework is proposed, which improves reasoning by transferring on-policy predictive signals from the final layer to intermediate representations and employs signed advantage-weighted Jensen--Shannon alignment to distill informative intermediate representations while preserving policy consistency under a unified acting policy.

Xin-Yu Liu, Darryl C. Jacob, Yang Zhou et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.