Skip to content

Category

natural language processing

2,926 papers

#natural language process... Conference Open access Aug 2026

Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts

Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.

Jaehee Kim, Jihoon Chung, Seoyoon Park et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Stratified Consistency Distillation for Natural Language Formalization

A fine-tuning-based Stratified Consistency Distillation approach that shows significant and consistent improvements in both Pass@K and the novel Equivalent Logical Similarity metrics, demonstrating the potential of advancing logical translation through consistency distillation.

Zhi-Chao Hou, Ferhat Erata, Joseph Lilien et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, raising questions about whether models truly follow the underlying logical structure. Studying this behavior is challenging because the symbolic components of logical problems, such as operators and predicates, are difficult to systematically manipulate in natural language. We introduce a tool-driven framework for generating controlled, label-preserving edits to logical reasoning problems. Our method operates on symbolic representations of first-order logic and constraint satisfaction problem tasks, enabling targeted modifications to logical operators and other structural components before translating them back into natural language. Using this framework, we evaluate various LLMs under cumulative and individual operator edits and analyze their behavior in response to these changes. Our quantitative and qualitative analyses show that LLM reasoning behavior under controlled operator edits is inconsistent, regardless of model size or family: models sometimes adapt correctly to structural changes but often fail to track their logical consequences. The results from this automated stress test enable an evaluation of language models across different dimensions and help measure the reliability of their reasoning.

Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi · 0 citations
#computer vision Preprint Aug 2026

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.

Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li et al. · 0 citations
#artificial intelligence Review Aug 2026

The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce

This work introduces the Differential Reasoning Router (DRR), a cost-aware framework for cold-start LLM annotation that jointly optimizes model selection and human escalation, enabling a gradual shift from human-heavy cold-start annotation toward high-confidence automated routing.

Cheng Lyu, Jingyu Zhang, Vinny DeGenova et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Label Semantic Expansion via Label Guided Neural Topic Modeling

A Label-Guided Neural Topic Model (LGNTM) is proposed, which learns dedicated label-aligned topics, grounds them in lexical and document semantic spaces, and preserves consistency between topic structures and label structures.

Hao-Jia Zheng, Yuyin Lu, Jun-Tian Huang et al. · 0 citations

When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection

Sarcasm detection is addressed through sarcasm detection, evaluating Qwen2.5-Omni and Qwen3-Omni on Mandarin Chinese and English under five modality conditions that decompose the contributions of lexical content, vocal semantics, and prosodic structure to reveal a shared stereotype of expressive prosody.

Yong-Jian Chen, Pengfei Wei, Yiqun Sun et al. · 0 citations
#natural language process... Preprint Aug 2026

When Errors Become Memories: Causal Pathway Tracing in Multi-Turn Memory-Augmented LLMs

A structural causal model (SCM)-based framework for cross-turn error propagation in memory-augmented LLMs is proposed, and experiments show that error influence generally decays with interaction distance, while the memory-update pathway contributes more persistent effects than question feedback.

Shu-Yao Xiao, Sheng-Ling Wang, Xuan Chen et al. · 0 citations
#natural language process... Preprint Aug 2026

ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives

ALTSTEER is an inference-time framework that couples selective intervention with refusal-anchored constructive redirection within a single inference pass, and uses an internal refusal-relevant signal to decide when to steer, and applies staged steering to shift generation from refusal-oriented control toward constructive alternatives.

Hoejoon Kwon, B. Lim, K. Kim et al. · 0 citations
#natural language process... Preprint Aug 2026

GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space

GPAgentBench-2K is introduced, the first Constrained MDP (CMDP) LLM-agent benchmark for primary-care clinical decision-making, constructed from expert-validated records of real-world GP encounters, and uncovers a clinical quality-safety gap.

Bo-Qi Chen, Xudong Liu, Y. Ao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation

This work proposes CPR (Critical-Point Routing), a token-level routing framework between a base model and its expert derivative, based on critical tokens where the base model fails but the expert succeeds, and achieves state-of-the-art across all settings.

Kwangmin Ki, Yunhun Nam, Jongheon Jeong et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.