Skip to content

Category

artificial intelligence

6,499 papers

#artificial intelligence Preprint Aug 2026

Open-World Semantic Segmentation with Sensitivity Modeling

This work addresses open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design.

Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki · 0 citations
#artificial intelligence Preprint Aug 2026

Scaling an Autoregressive Transformer for Single-Cell Generation

The first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model is found, finding the first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model.

A. Sharipov, Yusif Mukhtarov, Igor Molybog · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Training nGPT

The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. In this paper, we describe a practical training recipe for nGPT and evaluate it on modern hybrid Mamba-2--Transformer Mixture-of-Experts (MoE) models. The recipe introduces Logit Gradient Preconditioning, Logarithmic Learning Rate Decay, GatedAdamW, angular update control, and optional exploration mechanisms. Compared with an unnormalized model of the same hybrid MoE architecture trained with AdamW, the 30B-total-parameter nGPT model reaches the same validation loss using approximately half as many training tokens. The recipe scales across the models considered, which contain up to 30B total parameters.

Ilya Loshchilov, Boris Ginsburg · 0 citations

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

SABER-Math is introduced, the first fully automated benchmark for evaluating mathematical IR without expert annotation, and it is shown that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.

N. Georgiev, Maria Drencheva, Kseniia Ibragimova et al. · 0 citations
#artificial intelligence Preprint Jun 2026

AdaMem: Learning What to Remember with Adaptive Memory Policies for Personalized Agents

AdaMem is introduced, which uses adaptive natural-language Memory Policies to personalize what an agent writes to memory and demonstrates the promise of adaptive write control while exposing policy execution as a central limitation of current memory agents.

Xing-Yu Chen, Rui Wang, Zhaopeng Tu et al. · 1 citation · ⚡1

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning

The first systematic study of poisoning-based backdoor attacks on Speech Emotion Recognition systems with a focus on threats enabled by text-to-speech (TTS) generated audio is presented, revealing that TTS technology dramatically lowers the barrier to effective backdoor attacks.

Yong-Bin Huang, Xi-Hao Xie, Jia Zhang · 0 citations

Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication

Two recent studies \citep{jones2026llms, zeng2026lvlms} reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions. We control for task differences between the studies while directly comparing their prompting styles. We replicate the finding that models can coordinate efficient referring expressions when \textit{explicitly} prompted to do so, suggesting that other task differences are not responsible for divergent results. However, we also find that the same models fail to infer the need for communicative efficiency from a more \textit{implicit} prompt, highlighting critical differences between how humans and AI systems communicate.

Peter Zeng, Amie Paige, Wei-Ling Li et al. · 0 citations

Do Large Language Models Always Tell The Same Stories?

It is demonstrated that frontier models in particular converge on a "mean"generic narrative that approximates individual human stories but lacks the collective diversity of human authors, and it is shown that common mitigation strategies fail to meaningfully address this homogeneity.

K. ThennalD, Hans Ole Hatzel · 0 citations

Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens

This work proposes ASRD (Anchor Supervised Revocable Decoding), a training-free framework that operates within the embedding space and outperforms recent remasking baselines, achieving accuracy improvements of up to 6.4\% while accelerating inference throughput by up to 7.2$\times$.

Yizhen Yao, Qing-Lin Zhu, Runcong Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Emotional regulation improves deep learning-based image classification

Emotion significantly influences cognition, enhancing memory and learning under certain conditions. Drawing on this principle, emotion-augmented deep learning investigates how affective states can improve neural network architectures and learning paradigms, achieving better generalization than non-emotional models. However, existing methods often rely solely on objective neurophysiological factors, neglecting the role of subjectivity in emotion. To bridge this gap, the present study introduces Emotional Regulation, a novel framework for modeling emotion in deep learning through artificial subjective experience. The method employs pre-training based on affective stimuli, balancing non-emotional and emotionally-influenced responses in downstream task optimization. Extensive experimentation was conducted in image classification, pre-training ResNet and ViT architectures on four emotional datasets, using CIFAR-10 and -100 as target benchmarks. Results reveal improvements over the aforementioned backbones, providing evidence of Emotional Regulation as a promising method for defining emotion-augmented deep learning through artificial subjective experience. Furthermore, the proposed approach overcomes the related work in image classification based on CIFAR, revealing Emotional Regulation as the new state-of-the-art in emotion-augmented deep learning for large-scale vision datasets. The study also enforces evidence of the impact of affective states in improving machine learning tasks' optimization, encouraging further investigation on emotion-inspired architectures.

Riccardo Emanuele Landi, Jo\~ao M. F. Rodrigues, Marta Chinnici · 0 citations
#artificial intelligence Preprint Open access Sep 2026

WhiFlash: Accelerating Speculative Decoding with Token-Level Cross-Paradigm Routing

The autoregressive nature of large language models (LLMs) remains a significant bottleneck for inference, particularly in complex agentic workloads. While speculative decoding (SD) accelerates inference, current approaches rely on static drafting paradigms, utilising either autoregressive drafting models for reasoning or diffusion-based parallel drafting models for structured outputs. We empirically find that drafting accuracy fluctuates dramatically within a single sequence, leaving significant performance unrealised by static paradigms and coarse-grained routing. To address this volatility, we introduce WhiFlash, the first cross-paradigm SD method that unifies autoregressive and diffusion-based parallel drafting under a single token-level controller. WhiFlash adopts a fine-grained routing mechanism that employs either a lightweight entropy-based or a learned neural policy, both parametrised to provide a tunable balance between expected token gain and latency. To make high-frequency switching computationally viable, we introduce novel cache-management optimisations, Lazy Catch-up and KV-only Prefill, reducing switching overhead to below 7% of per-round latency. By capitalising on the complementary strengths of fundamentally distinct drafting architectures, WhiFlash achieves significantly higher acceptance lengths, yielding category-specific throughput gains of up to 69.6% over the state-of-the-art autoregressive EAGLE-3 and 37.3% over the diffusion-based DFlash.

Young D. Kwon, Miles Williams, Rui Li et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

TUX: Measuring Human--AI Tacit Understanding

As large language models (LLMs) increasingly act as collaborative partners, human--AI alignment is often evaluated through explicit task success, accuracy, or reward optimization. Yet many collaborative settings depend on tacit understanding: whether an agent can align with a human's evaluative stance or representational priors without clear objectives, communication, or feedback. To study this capacity, we develop a spectrum-placement task inspired by the social party game Wavelength, in which humans and agents independently place concepts along subjective spectra. We operationalize the Tacit Understanding Index (TUX) as a pairwise behavioral measure of similarity between human and agent judgments, and evaluate it with 241 human participants and 200 profile-conditioned LLM agents across four models. We find that nearest human--agent pairs in trait space achieve significantly higher TUX, suggesting that tacit alignment is associated with person-level characteristics rather than reflecting only random similarity. Regression analyses show that TUX becomes more explainable as predictor sets become richer, with individual traits, decision-making styles, and confidence improving over aggregate trait-distance baselines. These findings suggest that TUX provides a measurable behavioral signal of human--LLM tacit understanding, while revealing the limits of profile-based conditioning for capturing deeper representational alignment.

Yueshen Li, Hanyi Min, Vedant Das Swain et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.