Skip to content

Category

machine learning

4,920 papers

#machine learning Preprint Aug 2026

Selection-Aware Stress Testing for Interactive Agents

This work introduces Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks.

Yang Xu, Chenang Li, Jiefu Zhang et al. · 0 citations
#machine learning Preprint Aug 2026

S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation

S3C-LLM is introduced, a skill-guided and code-grounded agentic LLM for spectrum-to-structure elucidation that consistently outperforms current general LLMs and spectrum-specific models across spectra, while using less than 1/10th of SpectraLLM's training corpus.

Xuan-Le Zhao, Xinyu Cai, Xiang Cheng et al. · 0 citations
#machine learning Preprint Aug 2026

Fine-Tuning Low-Bit Models with Gradient in Quantized Code Space

This work proposes code surrogate gradient as the first order signal in deployable code space to acceleate optimization, and performs guided search to preserve deployment faithfulness in fine-tuning low-bit models across different quantization datatypes.

Shiguang Wu, Zhou-Chen Lin, Quan-Ming Yao · 0 citations
#machine learning Preprint Aug 2026

Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual Screening

This work validates the engineering feasibility of running industrial-scale trillion-parameter LLM-driven biomedical computing tasks on consumer hardware, establishing a new low-barrier paradigm for AI-powered early stage drug discovery.

Rui-Ya Xiao, Yi-Li Xu · 0 citations
#machine learning Preprint Aug 2026

What Emerges and What Breaks in Self-Play Driving

This work analyzes which traffic rules emerge from self-play and how closely they match human driving, and confirms that reward conditioning yields the intended diversity of driving behaviors.

Laur Sisask, Ardi Tampuu, Tambet Matiisen · 0 citations
#machine learning Preprint Aug 2026

Geometric Attractor Monitoring: A Robust and Frugal Framework for Multi-modal Industrial Robotic Cycles

This work reframe the monitoring problem through a framework based on Phase Space Reconstruction (PSR), and shows that aligning the algorithmic bias with the geometric properties of the target system yields a pragmatic, traceable and easily deployable approach perfectly tailored to the realities of industrial constraints.

Martin Bonsergent-Brachet, Jesse Read, D. Abboud · 0 citations
#machine learning Preprint Aug 2026

Reciprocity Separates Gradient Flow from Rotation in Conservative Physical Learning

Physical learning lets a trainable material or network use its own physical response to carry error signals, reducing the need for a separately programmed backward computation. We ask what determines whether such a system follows conventional gradient descent or evolves along a genuinely different learning trajectory. Our canonical model is a directed layered transport network in which every node redistributes a fixed amount of flow, so learning preserves positivity and total mass. In this model, conservation constrains only the allowable learning directions. Within the matched response class studied here, adjoint matching gives the physical output response a symmetric form. Non-negative mode-wise feedback then produces a reciprocal closed-loop response and a reweighted gradient flow. Adding an antisymmetric boundary component makes the closed-loop response rotational: the learning path can turn while the error driving that update still decreases at that moment. Turning is not automatically beneficial. Its finite-step effect is set by local curvature, and its accumulated effect also depends on step selection and on the new states visited along the path. Numerical consistency checks reproduce the exact response structure, predict the sign of the local effect across new network families, and show how trajectory drift can negate a local advantage. These results separate the roles of conservation, reciprocity, and nonreciprocity in physical learning.

Rui-Wu Niu, Xiao-Wen Bi, M. V. van Wyk · 1 citation
#machine learning Preprint Aug 2026

TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training

This work presents the first systematic characterization of SDC vulnerability across major computation interfaces in both the forward and backward passes of Transformer training, and proposes TrainSDC, a characterization-guided protection framework consisting of Q/K-path recomputation, residual-gain monitoring, and exponent-aware gradient scaling.

Zhijie Xia, Haotian Xu, Si-Yu Yun et al. · 0 citations
#machine learning Conference Open access Jun 2023

T3S: Improving Multi-Task Reinforcement Learning with Task-Specific Feature Selector and Scheduler

A novel MTRL framework called Task-Specific feature Selector and Scheduler (T3S) is proposed, which consists of two components: a feature selector and a task scheduler that consistently outperforms the state-of-the-art M TRL algorithms on various robotics manipulation tasks.

Yuan-Qiang Yu, Tianpei Yang, Yongliang Lv et al. · 4 citations
#machine learning Preprint Aug 2026

PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents

This work introduces PRACTICE, which trains a skill learner to discover and maintain a persistent skill library from past interaction trajectories while keeping the task executor frozen, and demonstrates that a compact skill learner delivers consistent performance improvements across successive library-update rounds for multiple frozen executors.

Ziyi Bai, Si-Qi Li, Ting-Lei Huang et al. · 0 citations
#machine learning Preprint Aug 2026

Do VLMs Share Safety Neurons Across Modalities?

A causal, neuron-level analysis of safety mechanisms in 10 VLMs, a two-stage detection pipeline with iterative ablation that accounts for self-repair, and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals are introduced.

Jia-Xuan Li, Jia-Hao Zhang, D. Vo et al. · 0 citations
#machine learning Preprint Aug 2026

TDDM-Melatt: A Decoupled Memory and Diffusion Framework for Generalizable Encrypted Traffic Classification

The proposed TDDM-Melatt, a disentangled memory-based traffic classification framework with diffusion-based data augmentation, provides a new and effective technical pathway for encrypted traffic classification in real-world network environments.

Zelang Chen, Qiming Yu, Zijia Song et al. · 0 citations

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.