Skip to content

Category

machine learning

5,133 papers

#machine learning Preprint Aug 2026

Fast Weight Attention for Continual Learning

This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.

Yi-Fan Zhang, Steve Ta, Jasper Zhang et al. · 0 citations
#machine learning Preprint Aug 2026

The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs

Overall, it is found that probing is an effective means to catch a range of different tool-calling errors, including errors arising from using an argument that has the wrong value but the correct type, which might not be recorded by standard logging frameworks.

Eric C. Yeats, Brendan Kennedy, Loc Truong et al. · 0 citations
#machine learning Preprint Aug 2026

Diffusion Distillation for Efficient Weather Ensembles

A supervised energy-distance distillation method is introduced that compresses a multi-step diffusion teacher into a single-step student by aligning student forecasts with teacher samples and ground-truth observations and preserves skill for extreme events.

Yiming Yang, Valentin Brekke, James Briant et al. · 0 citations
#machine learning Preprint Aug 2026

Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning

This paper considers the practical setting where the learner keeps a small proxy dataset, and proposes a dynamic, influence-aware client selection framework that estimates each client's potential utility to the learner's optimization objective using proxy influence signals on a learner-specific proxy set.

Yiming Xie, Linghui Su, Ningfang Mi · 0 citations
#machine learning Preprint Aug 2026

DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge

Results show that DART-FL dynamically adapts the inference-training resource split to time-varying inference demand and shifts the learning progress of high-demand tasks toward their burst periods, improving model accuracy when those tasks are frequently requested while maintaining comparable long-term multitask performance.

Yiming Xie, Pinrui Yu, Geng Yuan et al. · 0 citations
#machine learning Preprint Aug 2026

SafeStep: An Interactive Demonstration of Semantic Communication for Pedestrian Safety Monitoring

SafeStep is the first real-time semantic communication platform to make AoI-induced downstream degradation directly observable in live monitoring applications, and a recently proposed semantic communication design called Meta-VIB with five baseline transceivers.

Christian McDowell, Andrea Panebianco, Jeremiah Yang et al. · 0 citations
#machine learning Preprint Aug 2026

SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning

SegBench-GC is introduced, a controlled stress test of segmentation invariance that holds transitions, source trajectories, goal sampling, optimization settings, and evaluation fixed while varying only artificial backup boundaries and whether those boundaries retain continuation value.

Musa Shams · 0 citations
#machine learning Preprint Aug 2026

More Data Cannot Break a Symmetry: Identifiability by Design

Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus geometry bounds what any such alignment can identify, before data exist. The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling creates near-duplicates whose transposition is nearly free. We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention. In colour, where candidate geometries have closed form, we show that the failure is structural: sixty-four times the restart budget leaves a symmetric design unmoved while an asymmetric set at the same N recovers every time. Discriminating representational models and recovering a correspondence are essentially uncorrelated objectives (r = -0.02 over 3,000 subsets). Choosing nine colours by this diagnostic alone, without consulting any learned representation, moves all 93 model representations away from the degenerate point and cuts catastrophic alignment failures from 75% to 2% with the models, the layers, N and the solver all held fixed. The same risk arises wherever a regular design meets its candidate geometry's isometry group, including evenly spaced orientations, tones, or motion directions, and the check costs one function call before data collection.

Jing-Lin Xu, Christopher Kanan · 0 citations
#machine learning Preprint Aug 2026

Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics

An evolving benchmark suite of challenging, natively-spherical PDE datasets including a modified Galewsky jet, anomalous chained turbulence, Cahn-Hilliard decomposition, spherical Riemann shocks, Held-Suarez dry atmospheric transport and global ocean dynamics are released.

T. Muser, Giovanni Abati, Ivan Dokmanic · 0 citations
#machine learning Preprint Aug 2026

When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging

This work reveals that the recent Muon optimizer as a mechanism that regulates this factor by construction tightens the interference bound for both CL and MM, positioning Muon as a principled optimizer-centric approach complementary to existing solutions.

Shan Liu, Yuehan Yin, Yinghuan Shi et al. · 0 citations

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.