Skip to content

Category

machine learning

5,133 papers

#machine learning Preprint Aug 2026

A Deeper Analysis of Block-Sparse Featurizers

This work proposes several architectural changes to the BSF, including a Tournament Top-K selection rule that significantly reduces feature splitting, and extends the block paradigm to the crosscoder.

Alexandru-Iulius Jerpelea, Amith Ananthram · 0 citations
#artificial intelligence Preprint Aug 2026

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

This work introduces TraceML, which pairs human and agent work on the same competitions under one version-level schema, and releases the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.

J. Yan, Weiwei Sun, Si-Jie Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

On-policy Distillation with Verifiable Reward

This work reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, making it readily combinable with any policy gradient algorithm, such as GRPO.

Wenze Lin, Jiale Zhao, Xi-Tai Jiang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

A hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration that consistently improves the quality of both degraded and low-resolution images.

Saif Ahmed, Ashadullah Galib, S. R. R. Antu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

JuryProbe is introduced, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy, which estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift.

Tianxing Zhou, Ruixi Lin · 0 citations
#artificial intelligence Preprint Aug 2026

ED-CSP: Crystal Structure Prediction from Electron Diffraction

This work introduces ED-CSP, a machine learning framework that predicts crystal structures from chemical composition, atom count, and multiple detector-plane ED spot sets and establishes a benchmark for generative crystal structure prediction from sparse ED observations and provides a foundation for future transfer to experimental data.

Germain Poloudenny, Arnaud Demortière, Yael Fregier · 0 citations

Where Steering Signals Come From: Activation Source Selection in Activation Steering

Tail subtraction is introduced, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals, and suggests that steering depends on representations of what the model is about to do, not merely on what has already appeared.

Jiaran Ye, Lingxu Ran, Zijun Yao et al. · 2 citations
#artificial intelligence Preprint Jul 2026

On the Depth Scalability of Logic Gate Networks

Results indicate that scalable LGN depth requires both stable optimization and credit-preserving information access, and introduce Input-Anchored Logic Gate Networks (IALGN), in which each gate combines a private hidden spine with a direct input anchor.

Taegun An, Dohun Kim, Haebeom Lee et al. · 0 citations

The Discrete-Log Clock: How a Transformer Learns Modular Multiplication

The transformer reduces multiplication to addition in discrete-log space, implementing a "Discrete-Log Clock" algorithm analogous to Nanda et al.'s Clock algorithm for addition, which generalizes: matching the analysis basis to the algebraic structure of the task reveals interpretable structure where standard tools see noise.

Huu-Tuan Nguyen · 0 citations

TokenPilot: Cache-Efficient Context Management for LLM Agents

TokenPilot is presented, a dual-granularity context management framework that reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems.

Buqiang Xu, Z. Xue, Dian Chen et al. · 1 citation

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

LongDS is introduced, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget.

Kewei Xu, Xiaobe Lu, Shuofei Qiao et al. · 1 citation

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.