Skip to content

Category

natural language processing

2,394 papers

#artificial intelligence Preprint Jan 2026

To Retrieve or To Think? Cross-Boundary Context Evolution for Multi-hop Complex Reasoning

EvoCtx dynamically decides whether the next reasoning transition should cross the current evidence boundary through retrieval or refine the reasoning state within the existing context, and strategically alternates between boundary expansion and intra-boundary trajectory refinement.

Rubing Chen, Jian Wang, Wenjie Li et al. · 2 citations
#artificial intelligence Preprint Jan 2026

Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models

It is suggested that CoT prompting activates specific latent features to trigger reasoning, and that targeted intervention on these features offers an alternative pathway to elicit efficient reasoning behavior without explicit CoT prompting.

Zhenghao He, Guangzhi Xiong, Bohan Liu et al. · 6 citations · ⚡1

Kinship Data Benchmark for Multi-hop Reasoning

This work introduces KinshipQA, a benchmark designed to probe large language models' ability to perform multi-hop reasoning through reasoning over kinship relations, and demonstrates that KinshipQA yields a wide spread of outcomes and exposes systematic differences in multi-hop reasoning across models and cultural settings.

Tianda Sun, D. Kazakov · 1 citation

LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents

A novel architecture incorporating an explicit private working memory is proposed and it is demonstrated that this mechanism restores consistency with a fixed hidden state, establishing private state as a necessary component for PSIT-capable language agents.

Davide Baldelli, Alipanah Parviz, A. Zouaq et al. · 2 citations

Do Language Models Reason Across Languages?

This paper introduces a simple two-hop question answering setting, where answering a question requires making inferences over two multilingual documents, and finds that language models are more sensitive to language variation in answer-span documents than in those providing bridging information, despite the equal importance of both documents for answering a question.

Yan Meng, Wafaa Mohammed, C. Monz · 1 citation

Labels have Human Values: Value Calibration of Subjective Tasks

MultiCalibrated Subjective Task Learning (MC-STL), a framework that identifies latent value groups from annotations and enforces value-conditional calibration through value group-specific representations, is proposed and evaluated.

Mohammed Fayiz Parappan, Ricardo Henao · 1 citation

AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs

AdaFuse is an adaptive ensemble decoding framework that dynamically selects semantically appropriate fusion units during generation that establishes a synergistic interaction between adaptive ensembling and test-time scaling, where ensemble decisions guide targeted exploration, and the resulting diversity in turn strengthens ensemble quality.

Cheng Cui, Tianxin Wei, Ziyi Chen et al. · 6 citations · ⚡1
#natural language process... Preprint Jan 2026

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

Field-Aware Contrastive Decoding (FACD) is proposed, a training-free strategy that amplifies suppressed disposition-sensitive signals, significantly closing the performance gap without sacrificing moral-character performance.

Yonghyun Jun, Junhyuk Choi, Jihyeon Park et al. · 1 citation
#artificial intelligence Preprint Jan 2026

EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning

EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction for epidemiological question answering over research literature, comprising three subsets built from open-access articles across diverse diseases.

Mingyang Wei, De-Hai Min, Zewen Liu et al. · 0 citations

DIP: Dynamic In-Context Planner For Diffusion Language Models

DIP is proposed, a context-optimization algorithm based on average verified confidence that dynamically ranks and inserts in-context examples during generation, rather than providing all examples up front.

Yang Li, Han Meng, Chenan Wang et al. · 1 citation

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.