Skip to content

Category

natural language processing

2,394 papers

Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference, and enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity.

Yu Xia, Zhouhang Xie, Xin Xu et al. · 0 citations

Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?

This work proposes a pipeline for automatically generating step-by-step linguistic reasoning traces from Universal Dependencies treebanks, dictionaries, and grammar-rule banks and shows that linguistic reasoning traces are most effective as inference-time guidance in ICL, which substantially improve translation performance across models, languages, and metrics.

Renhao Pei, Yihong Liu, Sampo Pyysalo et al. · 0 citations
#artificial intelligence Review Jun 2026

The Unsampled Truth: Quantifying Prompt Artifacts in LM Psychometrics

When prompting language models for psychometric assessment, researchers assume that the responses reflect the injected persona and the meaning of the survey item. We test this premise using a diagnostic design that crosses five semantically distinct baseline personas with five semantically equivalent variants of each of four prompt components (persona wording, task instruction, item wording, option symbol). Measuring the 1-Wasserstein distance between the resulting response distributions and partitioning the variation among the five components allows for the separation of target effects from prompt artifacts. We apply the framework to 13 open-weight small language models (0.6B to 14B) on the Big Five Inventory and the Short Dark Triad. We find that in most models, the task instruction and option symbol displace response distributions further than paraphrasing the persona description or the item itself. For a substantial share of items, the artifact share of explained variation exceeds 50%; non-semantic changes of the prompt account for more response variation than the baseline personas. Our framework lets researchers quantify these prompt artifacts before interpreting psychometric output.

Nils Schwager, Christoph Hau, Simon Münker et al. · 0 citations
#computer vision Jun 2026

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

It is shown that a lightweight LLM-only DPO update on tiny single-object-pair synthetic data mitigates the bias, lifting four-way robust accuracy by up to 100 points on synthetic data, and by 68.1 points on broader evaluation datasets WhatsUp, SpatialMQA-Direct, and VSR.

Chuan Ma, Qianying Liu, Tomoyuki Obuchi et al. · 0 citations

Linguistics-Aware Non-Distortionary LLM Watermarking

LUNA is a linguistically adaptive watermark that combines model-free detection with single-token non-distortion under the standard random-key model, and is the only method that simultaneously achieves AUROC>0.99 and an absolute median perplexity shift below 0.1.

Shinwoo Park, Hyejin Park, Hyeseon An et al. · 0 citations

Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning

SelSkill is proposed, a dual-granularity preference-learning framework for selective skill invocation that formulates skill use as a skill-or-skip decision, uses predictive uncertainty to prioritize candidate decision points, and constructs controlled invoke-skip preference pairs from shared trajectory prefixes.

Chishui Chen, Jiaye Lin, Te Sun et al. · 1 citation · ⚡1
#natural language process... Preprint May 2026

Unlocking Fine-Grained Translation Quality Estimation in LRMs through Mutually Boosting Implicit and Explicit Reasoning

This paper proposes a simple two-stage training framework that enables the mutually boosting of implicit (layer-wise) and explicit (token-wise) reasoning capabilities, and provides evidence for the mutually boosting between implicit and explicit reasoning.

R. Dang, X. Wang, Zhejian Lai et al. · 0 citations

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities. To address this, we introduce FBHM, a systematically curated benchmark of Functionality Based Hateful Memes constructed along two orthogonal axes: 25 distinct rhetorical functionalities and 10 target communities (5,000 memes total). Benchmarking state-of-the-art VLMs reveals a severe generalization gap: models highly accurate on standard datasets catastrophically drop to near-random performance on FBHM, proving they exploit dataset-specific heuristics rather than robust multimodal reasoning. To efficiently close this gap, we propose LSV (learnable steering vectors), an ultra-low data regime strategy that applies a causal intervention objective on as few as 500 steering samples (50 unique base memes), boosting FBHM performance by ~30 Macro-F1 points while outperforming in-context learning and PEFT without degrading source-domain performance.

Paramananda Bhaskar, Naquee Rizwan, Daksh Jogchand et al. · 0 citations
#natural language process... Preprint May 2026

Reinforcement Learning Can Amplify Emergent Misalignment from Harmless Rewards

It is shown that rewarding narrow, overtly misaligned behavior produces substantially higher general-domain misalignment than sample-matched SFT, and that EM from RL can be induced by reward signals that could plausibly arise naturally, such as unpopular aesthetic preferences or poor rhetorical appeals.

Magnus Jørgenvåg, David Kaczér, Lasse Ruttert et al. · 2 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.