Skip to content

Category

machine learning

4,920 papers

Divide-and-Conquer Modeling for the CTF-4-Science Lorenz Benchmark

This submission documents the divide-and-conquer modeling strategy developed for the CTF-4-Science Lorenz Chaotic Systems Challenge at AI-DEEDS 2026, which shows that bounded, scenario-specific updates can outperform broad model replacement on mixed chaotic forecasting benchmarks.

Shun-Dong Li · 0 citations
#machine learning Preprint Jun 2026

QueryGraph: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation

This work introduces a system that converts natural language queries into structured graphs and executes them via a deterministic planner, which uses depth-first search to resolve dependencies and combine results across tools, improving reliability and enabling queries beyond traditional keyword-based search.

A. Chakravarthy, Vidhi Kulkarni, Duen Horng Chau · 0 citations

TriHead-GAN: A Generative Adversarial Network with Triple-Head Discriminator for Carbon Emission Time Series Generation

TriHead-GAN is proposed, a Transformer-based adversarial framework whose triple-head discriminator jointly supervises three complementary aspects of the joint distribution: distributional authenticity via a Wasserstein critic, cross-variable dependency via leakage-free regression of the target variable, and step-wise temporal smoothness via adjacent-difference prediction.

Ze-Sen Wang, Lijuan Lan, Yong-Gang Li et al. · 0 citations

GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution

This work reframe attribution as subset-level counterfactual utility prediction and introduces GRASP, an interaction-aware surrogate that more than doubles the task-level rank correlation for counterfactual subset fidelity while reducing upfront artifact construction costs by nearly an order of magnitude.

Yue Min, Rui Chen, Yujun Li · 0 citations

Mamba-Assisted Non-Markovian Closure for Reduced-Order Modeling

The Mamba-Assisted Closure (MAC) framework is proposed, which employs a Mamba-based sequence model to predict the closure from the resolved trajectory and couples the learned closure with the reduced-order governing equations through a numerical integrator to advance the resolved variables in time.

Zhifei Wei, S. Qadeer, Panos Stinis · 0 citations

GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning

GRZO is a Group-Relative Zeroth-Order optimizer that draws one pseudo-independent perturbation per mini-batch example and aggregates the per-example losses through group-relative normalization, raising the effective gradient-direction count from one to the batch size at no additional forward cost while preserving inference-level memory.

L. Tan, Yequan Zhao, Yifan Yang et al. · 0 citations

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

This work proposes a novel block-diagonal Riemannian metric derived from the pullback of the Frobenius inner product and develops a Riemannian gradient descent algorithm that uses a tuning-free Gaussian step size and scales linearly in the number of observed entries per iteration.

Pratik Jawanpuria, Ankish Chandresh, Bamdev Mishra · 0 citations

LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training

This work proposes LaRA, a layer-wise representation analysis framework for detecting contamination in RL post-trained LLMs and finds that contamination produces progressive geometric deviations across layers, including amplified perturbation sensitivity, stronger directional collapse, and enhanced local rigidity.

Minju Gwak, Minseok Kwak, Dongseok Lee et al. · 0 citations

OISD: On-Policy Internal Self-Distillation of Language Models

The OISD framework is proposed, which improves reasoning by transferring on-policy predictive signals from the final layer to intermediate representations and employs signed advantage-weighted Jensen--Shannon alignment to distill informative intermediate representations while preserving policy consistency under a unified acting policy.

Xin-Yu Liu, Darryl C. Jacob, Yang Zhou et al. · 0 citations

Locality-Aware Redundancy Pruning for LLM Depth Compression

It is shown that inter-layer redundancy can be either localized or globally distributed depending on the LLM architecture, and Representation Locality Score (RLS) is introduced, derived from global inter-layer hidden-state similarity.

Vincent-Daniel Yun, Youngrae Kim, Woosang Lim et al. · 1 citation · ⚡1

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.