Skip to content

Category

machine learning

6,259 papers

Not All LLM Reasoning is Visible in the Chain-of-Thought

This work demonstrates a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks and indicates that frontier models already perform consequential computation with no interpretable trace in their output tokens.

Vatsal Baherwani, Tom Goldstein, Ashwinee Panda · 4 citations · ⚡2
#machine learning Preprint Aug 2026

LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction

LM-X is introduced, which organizes prediction across task, event, and motor scales without claiming anatomical correspondence and shows that explicit multi-timescale predictive state can strengthen control while exposing interpretable internal estimates.

Jin Lou, Zhi Jing, Xu-Peng Wang et al. · 0 citations
#machine learning Preprint Aug 2026

Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity

A unified framework, KENDO (Kernel ENsemble Disagreement-aware Operator), is proposed that integrates Ensemble Gaussian Processes (EGP) with disagreement-aware acquisition strategies and extends the approach to multi-objective optimization via random scalarization that preserves the single-optimizer conditioning structure.

Heng Zhang, Hao-Tian Xiang, Konstantinos D. Polyzos et al. · 1 citation
#machine learning Preprint Jul 2026

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

These results establish a world-model design for QAS whose value lies in decision-useful feedback rather than exact energy prediction, and establish a world-model design for QAS whose value lies in decision-useful feedback rather than exact energy prediction.

Jiayang Niu, Yan Wang, Jie Li et al. · 0 citations
#machine learning Conference Jul 2026

Adaptive Confidence-Weighted Expansion for Trustworthy Multi-omics Multimodal Fusion

Adaptive Confidence-weighted Expansion (ACE), a novel framework to enhance the trustworthiness of multimodal fusion models, provides a more stable and robust data fusion method that facilitates the use of multimodal learning in addressing high-stakes problems.

Mohammad Raahemi, Ali Sekhavati, Alireza Maleki et al. · 0 citations
#machine learning Preprint Aug 2026

What Neural Network Field Theory Can and Cannot Realise on a Computer

A no-go theorem is used to separate four versions of neural network field theory, according to whether the defining object is the finite width ensemble or its infinite width limit, and whether the target the authors want to compute is a quantum or an effective field theory.

Thomas R. Harvey · 1 citation

Robust Chance-Constrained Optimization using a Continuous Parameter Space Wasserstein-2 Ambiguity Set of Gaussian Mixtures

A novel formulation of a Wasserstein-2 metric that uses the Bures-Wasserstein (BW) metric over probability measures with finite second moments is developed, which allows the worst-case distribution to endogenously determine both how many mixture components receive mass and where their means and covariances lie within a continuous support.

Shibshankar Dey, Sanjay Mehrotra · 0 citations
#machine learning Preprint Jul 2026

An End-to-End Hybrid Quantum--Classical Sampling Workflow for Discrete Markov Random Fields: A Reproducible Case Study

Sampling from discrete Markov random fields (MRFs) is a hard problem and amplitude-encoded i.i.d. sampling for small MRFs where $2^n$ target probabilities are precomputed classically is studied to allow a clean comparison against classical MCMC based on independent circuit samples.

A. Mazumder · 0 citations
#machine learning Preprint Jun 2026

Closing the Operational Gap in Semantic Caching

Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usable at a fixed threshold. We show this mismatch leads to systematically poor deployment choices, as models with the highest PR-AUC are often the worst in operation. We introduce Precision--Cache Hit Ratio (P-CHR) AUC, a cache-aware metric that measures precision across cache utilization levels, and Operational Retention Rate (ORR), which captures how much offline ranking quality survives at deployment. We decompose the operational gap between offline and deployed quality into a recoverable threshold-utility component and an irreducible structural component fixed by the dataset's positive rate. Our experiments show that the threshold-utility gap is governed by the training objective rather than data scale, and yields only to re-normalizing scores over the candidate pool or changing the training objective. Ultimately, model selection for semantic caching is a threshold-utility problem, not a ranking one, and measuring it is the first step to closing the gap.

Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev et al. · 1 citation

Online Learning-to-Defer with Varying Experts

An online multiclass L2D algorithm that combines queried-action bandit feedback with a dynamically varying pool of experts is introduced that achieves expected true-deferral regret under a concentrated-score condition.

Duy Hoang Dang, Yannis Montreuil, Maxime Meyer et al. · 5 citations

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

DiffAnon is proposed, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation, and is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control.

Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews et al. · 0 citations

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.