Skip to content

Category

machine learning

5,133 papers

#machine learning Preprint Jun 2026

Closing the Operational Gap in Semantic Caching

Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usable at a fixed threshold. We show this mismatch leads to systematically poor deployment choices, as models with the highest PR-AUC are often the worst in operation. We introduce Precision--Cache Hit Ratio (P-CHR) AUC, a cache-aware metric that measures precision across cache utilization levels, and Operational Retention Rate (ORR), which captures how much offline ranking quality survives at deployment. We decompose the operational gap between offline and deployed quality into a recoverable threshold-utility component and an irreducible structural component fixed by the dataset's positive rate. Our experiments show that the threshold-utility gap is governed by the training objective rather than data scale, and yields only to re-normalizing scores over the candidate pool or changing the training objective. Ultimately, model selection for semantic caching is a threshold-utility problem, not a ranking one, and measuring it is the first step to closing the gap.

Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev et al. · 1 citation

Online Learning-to-Defer with Varying Experts

An online multiclass L2D algorithm that combines queried-action bandit feedback with a dynamically varying pool of experts is introduced that achieves expected true-deferral regret under a concentrated-score condition.

Duy Hoang Dang, Yannis Montreuil, Maxime Meyer et al. · 5 citations

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

DiffAnon is proposed, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation, and is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control.

Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews et al. · 0 citations

Deflation-PINNs: Learning Multiple Solutions for PDEs and Landau-de Gennes

The results show that Deflation-PINNs can successfully identify and characterize multiple distinct crystal structures: a single unsupervised run recovers all six stable states of the benchmark and the discovered branches are refined to percent-level accuracy.

Sean Disarò, R. Maity, Aras Bacho · 0 citations
#machine learning Preprint Mar 2026

Agentic-Kube: A Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-Objective Kubernetes Scheduling

The architecture decomposes multi-objective scheduling into a tripartite optimisation space managed by dedicated sub-agents for cost minimisation, anti-affinity fault tolerance, and vector resource balancing, and Agentic-Kube consistently achieves Pareto-efficient placements.

Hamed Hamzeh · 0 citations

FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation

The results clearly demonstrate that FlowCorrect learns from very few demonstrations and enables fast, sample-efficient, incremental, human-in-the-loop corrections of generative visuomotor policies at deployment time in real-world robotics.

Edgar Welte, Yitian Shi, R. Wolf et al. · 4 citations

Robust Assortment Optimization from Observational Data

This work uncover and identify the notion of ``robust item-wise coverage''as the minimal data requirement to enable sample-efficient robust assortment learning and bridges the gap between robustness and statistical efficiency in assortment learning.

Miao Lu, Yuxuan Han, Han Zhong et al. · 0 citations

Learning Fast Monomial Orders for Gröbner Basis Computations

The resulting learned policies consistently outperform standard heuristics and resist distillation into simple interpretable models, providing empirical evidence that deep reinforcement learning allows the agents to exploit non-linear geometric structure beyond the scope of traditional heuristics.

R. Bunch, A. Ergür, Melika Golestani et al. · 0 citations

Prequential posteriors

This work introduces prequential posteriors, based upon a predictive-sequential (prequential) loss function, and proves that, under mild conditions, both the prequential loss minimizer and the prequential posterior concentrate around parameters with optimal predictive performance.

S. Roy, R. Everitt, Christian P. Robert et al. · 0 citations

Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity

Depression symptom severity was associated with five lexical features, including reductions in word count measures, use of first-person plural pronouns and positive word frequency, andLexical features and vector embeddings did improve prediction accuracy beyond baseline models.

A. Tokareva, J. Dineley, Z. Firth et al. · 0 citations

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.