Key contributions include proposing a multi-task learning framework for jointly optimizing visual quality and emotion, establishing the inaugural VAWE-Art dataset comprising 5,000 AI-generated images with 20-dimensional emotional annotations, and providing computational foundations for emotion-controllable generative art systems.
Hengju Gang· Journal of Engineering, Proj...· 0 citations
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.
Shu-Yi Fan, Boyuan Deng, Mengyu Xu et al.· 0 citations
This paper introduces a lightweight probe for multi-step gradients of pretrained weights that incurs no additional GPU memory cost and only marginal time overhead, and employs a spectrum-aware, importance-based rank allocation and optimal initialization derived from multi-step gradients.
Experimental evaluations demonstrate that the PTO framework enhances dialogue agents' performance in goal-oriented conversations within the domain of Motivational Interviewing, and incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies.
Lior Baruch, Moshe Butman, K. Bar et al.· 2 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work addresses open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design.
Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki· 0 citations
The first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model is found, finding the first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model.
A. Sharipov, Yusif Mukhtarov, Igor Molybog· 0 citations
A practical training recipe for the normalized Transformer and its evaluation on modern hybrid Mamba-2--Transformer Mixture-of-Experts models shows that the 30B-total-parameter nGPT model reaches the same validation loss using approximately half as many training tokens.
SABER-Math is introduced, the first fully automated benchmark for evaluating mathematical IR without expert annotation, and it is shown that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.
N. Georgiev, Maria Drencheva, Kseniia Ibragimova et al.· arXiv.org· 0 citations
AdaMem is introduced, which uses adaptive natural-language Memory Policies to personalize what an agent writes to memory and demonstrates the promise of adaptive write control while exposing policy execution as a central limitation of current memory agents.
Xing-Yu Chen, Rui Wang, Zhaopeng Tu et al.· 1 citation· ⚡1
The first systematic study of poisoning-based backdoor attacks on Speech Emotion Recognition systems with a focus on threats enabled by text-to-speech (TTS) generated audio is presented, revealing that TTS technology dramatically lowers the barrier to effective backdoor attacks.
Two recent studies \citep{jones2026llms, zeng2026lvlms} reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions. We control for task differences between the studies while directly comparing their prompting styles. We replicate the finding that models can coordinate efficient referring expressions when \textit{explicitly} prompted to do so, suggesting that other task differences are not responsible for divergent results. However, we also find that the same models fail to infer the need for communicative efficiency from a more \textit{implicit} prompt, highlighting critical differences between how humans and AI systems communicate.
Peter Zeng, Amie Paige, Wei-Ling Li et al.· arXiv.org· 0 citations
It is demonstrated that frontier models in particular converge on a "mean"generic narrative that approximates individual human stories but lacks the collective diversity of human authors, and it is shown that common mitigation strategies fail to meaningfully address this homogeneity.
K. ThennalD, Hans Ole Hatzel· arXiv.org· 0 citations