Skip to content

Category

natural language processing

3,231 papers

The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior

The stability of safety refusal decisions across random seeds and temperature settings is investigated by investigating the stability of safety refusal decisions across random seeds and temperature settings to demonstrate that single-shot safety evaluations are insufficient for reliable safety assessment and that evaluation protocols must account for stochastic variation in model behavior.

E. Larsen · 4 citations
#artificial intelligence Preprint Nov 2025

Think-at-Hard: Dynamic Looped Transformers for Improved Reasoning

This work proposes Think-at-Hard (TaH), a looped transformer optimized for selective iteration that employs a lightweight neural decider to trigger latent iteration, only at tokens likely to be incorrect after the standard forward pass.

Tianyu Fu, Yichen You, Ze-Kai Chen et al. · 0 citations

PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering

PRISM, an agentic retrieval framework that leverages large language models in a structured loop to retrieve relevant evidence with high precision and recall, achieves higher retrieval accuracy while filtering out distracting content, enabling downstream QA models to surpass full-context answer accuracy while relying on significantly less irrelevant information.

Md Mahadi Hasan Nahid, Davood Rafiei · 8 citations
#artificial intelligence Preprint Sep 2025

OceanGym: A Benchmark Environment for Underwater Embodied Agents

OceanGym is introduced, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments, and reveals substantial gaps between state-of-the-art MLLM-driven agents and human experts.

Yida Xue, Mingjun Mao, Xiangyuan Ru et al. · 0 citations

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

Safety-aware Contrastive Decoding (SafeCoDe) is introduced, a lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context that consistently improves context-sensitive refusal behaviors while preserving model helpfulness.

Zheyuan Liu, Zhangchen Xu, Guangyao Dou et al. · 6 citations

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

This work introduces a 98% automated pipeline to produce high-quality Quranic datasets and presents a novel ASR-based approach for pronunciation error detection utilizing the authors' custom Quran Phonetic Script (QPS) to encode Tajweed rules (unlike the IPA standard for Modern Standard Arabic).

Abdullah Abdelfattah, Mahmoud I. Khalil, Hazem M. Abbas · 1 citation

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

This work introduces a controlled setting to study the causes and training dynamics of cross-lingual knowledge transfer by training small Transformer models from scratch on synthetic multilingual datasets and suggests methods to encourage representational unification as part of training that would improve LLMs'cross-lingual transfer.

C. Blum, Katja Filipova, Ann Yuan et al. · 5 citations
#artificial intelligence Preprint Jul 2025

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

Cognitive Chain-of-Thought (CoCoT) is introduced, a reasoning framework that structures vision-language-model reasoning through three cognitively inspired stages: Perception, Situation, and Norm, showing that structuring model reasoning through cognitively grounded stages enhances interpretability and social alignment, laying the groundwork for more reliable multimodal systems.

Eunkyu Park, Wesley Hanwen Deng, Gunhee Kim et al. · 3 citations

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4

Long Story Short: Story-level Video Understanding from 20K Short Films

This work proposes Short-Films 20K (SF20K), the largest publicly available movie dataset, and accompanies this dataset with SF20K-Test, a manual, open-ended question answering benchmark, showing that instruction tuning on the large-scale dataset substantially improves model performance.

Ridouane Ghermi, Xi Wang, Vicky Kalogeiton et al. · 11 citations · ⚡1

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.