Skip to content

Category

natural language processing

2,926 papers

#machine learning Preprint Aug 2026

Kathleen Remembers: Length-Invariant One-Shot Recall Without Attention

This work adds to the Kathleen trunk a second memory layer -- a"notebook": a fixed-key holographic (HRR) associative store with a learned local write gate, a self-gating raw read, and write-triggered forgetting -- 25K parameters that attach to the logits of any trunk.

George Fountzoulas · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering

CREST is proposed, an inference-time alignment method that steers base model hidden representations using safety directions extracted from a guidance model of any family, avoiding token-level structural limitations entirely and outperforming baselines by up to 22.2\% on safety benchmarks.

Jin Gan, Xin Li, Jun Luo · 0 citations
#artificial intelligence Preprint Aug 2026

Using Prosody to Predict Syntactic Structure

This work quantifies the interaction between prosodic features and syntactic representations as their mutual information, and provides a general-purpose framework for estimating this quantity over large speech-text corpora using multimodal language models.

Junghyun Min, Alex Warstadt, Tamar I. Regev et al. · 0 citations
#artificial intelligence Preprint Aug 2026

VIBE: Video Instruction-aligned Background music gEneration

VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjective qualities with a structured 5-stage training curriculum.

Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj et al. · 0 citations
#machine learning Preprint Aug 2026

Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment

Evolutionary Soups, a mixture-of-experts framework for fine-grained generation control, with gating networks trained via an evolutionary algorithm, achieves the best hypervolume, linear utility, and Tchebyshev utility among controllable methods on all tasks.

Lingxiao Kong, Steffen Staab, Cong Yang et al. · 0 citations
#machine learning Preprint Aug 2026

When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs

This work presents a systematic mechanistic analysis of multilingual safety using sparse autoencoder features, sparse interpretable directions in the residual stream associated with harmful and harmless model behavior across three instruction-tuned LLMs, eight languages, and all model layers to qualify the language-universality of safety alignment as architecture-dependent and offer a mechanistic account of multilingual safety interventions.

Apoorva Upadhyaya, Sandipan Sikdar · 0 citations
#machine learning Preprint Aug 2026

Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer Evidence

This is the first work to formulate compression-aware abstention as a learning problem, in which a model learns to answer when supporting evidence survives compression and abstain when it does not, and controlled-deletion experiments show that the learned behavior is driven by evidence content rather than input length alone.

Mohammadali Khodabandehlou, Bhaskar Krishnamachari · 0 citations
#machine learning Preprint Aug 2026

Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation

Experiments show Influence-Directed Adaptive On-Policy Distillation (IDA-OPD), rather than relying on costly full-vocabulary Forward-KL objectives, preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability.

Run Yang, Runpeng Dai, Jie Sun et al. · 0 citations
#machine learning Preprint Aug 2026

Cross-lingual Functional Vectors for Emotion Detection in Large Language Models

This work examines whether FVs extracted from a source language can steer task behavior in another language under both standard clean and perturbed zero-shot settings without providing demonstrations during inference, and observes that each LLM exhibits a relatively stable optimal range of attention heads for constructing effective FVs, and the pattern remains consistent across languages.

Jieying Xue, Phuong Minh Nguyen, Minh Le Nguyen et al. · 0 citations
#machine learning Preprint Aug 2026

TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimization

TACS is proposed, a trajectory-aware candidate selection framework for jailbreak suffix optimization that augments per-step evaluation with a trajectory-aware proxy and stabilizes selection with reference-policy regularization and a discriminator-estimated chi-squared correction, encouraging choices that remain effective beyond the current step.

Shi-Liang Xiao · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.