Skip to content

Category

machine learning

8,448 papers

#artificial intelligence Preprint Open access Sep 2026

Learning to Construct Practical Agentic Systems

Automated design and optimization of agentic LLM-based systems leads to sophisticated systems that substantially improve result quality over off-the-shelf agentic patterns. However, studies of fielded agentic systems show that production systems focus much more on issues such as simplicity, controllability, and predict...

Aditya Kumar, Zhihan Lei, Jerry Yan et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Linear Separability of Activation Representations after Supervised Fine-Tuning on Incorrect Responses: A Study of Synthetic Dishonesty in Large Language Models

When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverable trace in its internal activations? We study this question in a controlled model-organism setting using five transformer architectures spanning 1.4 to 9 billion parameters. For e...

Vahideh Zolfaghari · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Gradient-Free Training of Spiking Neural Networks via Low-Rank Evolution Strategies

Spiking Neural Networks (SNNs) offer compelling energy efficiency on neuromorphic hardware, yet their training remains challenging because the discrete spike threshold is non-differentiable. Surrogate-gradient methods sidestep this by approximating the derivative, but they impose backpropagation infrastructure that is...

Dhruv Patankar, Sachit Ramesha Gowda · 0 citations
#artificial intelligence Preprint Open access Sep 2026

The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer

This book provides a compact, derivation-oriented introduction to the mathematical foundations of modern generative artificial intelligence. Rather than surveying every recent architecture or implementation detail, it develops a coherent route through the ideas connecting major families of generative models, from PCA,...

Tianhua Chen · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation

Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling K responses yields a belief state - responses a model deems plausible. Existing work exploits this representation for narrow tasks like either decoding or selective prediction, an...

Joris Baan, Wilker Aziz, Barbara Plank et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts

Detecting Schwartz values in political texts is hard: cues are often implicit, and neighboring values differ by fine distinctions. Two remedies are widely assumed to help: more surrounding document text, and explicit moral knowledge. Knowledge-based retrieval has improved benchmarks elsewhere, but whether either transf...

V\'ictor Yeste, Paolo Rosso · 0 citations
#artificial intelligence Preprint Open access Sep 2026

EmoTrack: Clinical-Semantic Modeling for Text-Based Depression Severity Estimation

Text-based counseling provides a valuable source of information for assessing depression severity. We study prediction of the total score on the eight-item Patient Health Questionnaire (PHQ-8), a self-report measure of depression severity, from counseling transcripts. Clinical-based methods rely mainly on large languag...

Zhaomin Wu, Jiayi Li, Bingsheng He · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models

Neuro-Symbolic Concept-based Models (NeSy-CBMs) are a family of architectures that integrate neural networks with symbolic reasoning for enhanced reliability in high-stakes applications. They work by first extracting high-level concepts from the input and then inferring a task label from these compatibly with given log...

Samuele Bortolotti, Emanuele Marconato, Andrea Pugnana et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound regime does not mean the model has fully utilized its organic corpus. In this paper, we introduce SynPro, a synthetic data generation framewo...

Zichun Yu, Chenyan Xiong · 0 citations
#artificial intelligence Preprint Open access Sep 2026

EmoMind: Decoding Affective Captions from Human Brain fMRI

Decoding visual experience from brain activity has advanced substantially, but current brain-to-text systems largely recover semantic content while discarding affect. Additionally, language models can generate emotional text when prompted with categorical labels, but such labels collapse rich inter-subject variability...

Bilal A. Mohammed, Lin Gu, Ruogu Fang · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring w...

Isabella Luong, Joyee Chen, Sankalpa Ghose et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

Video world models should predict future appearance in a way that remains consistent with 3D scene structure, camera motion, and lens geometry. Existing attention-level camera encodings, however, either describe each token only by its viewing ray---without locating scene content along that ray---or assume pinhole proje...

Seonghyun Jin, Youngmin Kim, Sunwoo Park et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.