Skip to content

Category

natural language processing

2,394 papers

#natural language process... Preprint Mar 2026

SafeMath: Safe Solutions for Unsafe Math Word Problems

Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we study an underexplored issue concerning harmful and toxic mathematical word problems. We show that math questions, particularly those framed as natural language narratives, can serve as a subtle medium for propagating biased, unethical, or psychologically harmful content, with heightened risks in educational settings involving children. To support a systematic study of this phenomenon, we introduce ToxicGSM, a dataset of 1.9k arithmetic problems in which harmful or sensitive context is embedded while preserving mathematically well-defined reasoning tasks. Using this dataset, we audit the behaviour of existing LLMs and analyse the trade-offs between safety enforcement and mathematical correctness. We further propose SafeMath -- a safety alignment technique that reduces harmful outputs while maintaining, and in some cases improving, mathematical reasoning performance. Our results highlight the importance of disentangling linguistic harm from math reasoning and demonstrate that effective safety alignment need not come at the cost of accuracy.

Sagnik Basu, Subhrajit Mitra, Aman Juneja et al. · 0 citations

OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models

OmniACBench, a benchmark for evaluating context-grounded acoustic control in omni-modal models, is introduced and three common failure modes are identified-weak direct control, failed implicit inference, and failed multimodal grounding-providing insights for developing models that can verbalize responses effectively.

Seunghee Kim, B. Park, Kyudan Jung et al. · 1 citation

MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment

This work introduces MemoNoveltyAgent, a multi-agent system designed to generate comprehensive and faithful novelty reports, and proposes a RAG-augmented checklist evaluation method that enables reliable and evidence-grounded assessments.

Jiajun Hou, Hexuan Deng, Wenxiang Jiao et al. · 0 citations

Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm

This study tested five Large Language Models and compared their performance to that of human controls using an adapted version of a text-based tool widely used in human ToM research, revealing a performance gap between the models.

Anna Babarczy, András Lukács, Péter Vedres et al. · 1 citation
#natural language process... Preprint Mar 2026

PACE-RAG: Patient-Aware Contextual and Evidence-Constrained RAG for Clinical Drug Recommendation

Patient-Aware Contextual and Evidence-Constrained RAG personalizes recommendations by first extracting patient-specific clinical features, retrieving cases around these features, and then refining the final prescription using the patient's current symptoms, active medication history, and focus-specific prescribing tendencies.

C. Huh, Hyunmin Hwang, J. Shin et al. · 0 citations

QAQ: Bidirectional Semantic Coherence for Selecting High-Quality Synthetic Code Instructions

QAQ, a novel data selection framework that evaluates data quality from the reverse direction: how well can the answer predict the query conditioned on the query, is proposed and Reverse Mutual Information (RMI) is defined to quantify the information gain about the query conditioned on the answer.

Jiayi Lei, Mingqing Ma, Yu Duan et al. · 0 citations
#natural language process... Preprint Feb 2026

QQ: A Language Metadata Toolkit for Multilingual NLP

This work presents QQ, a metadata toolkit and browser explorer that compiles language metadata sources into a graph of language varieties, scripts, regions, identifiers, names, and relations, and exposes it through a Python API, a command-line interface, and a browser-based explorer.

Wessel Poelman, Yiyi Chen, Miryam de Lhoneux · 1 citation

Small Reward Models via Backward Inference

FLIP (FLipped Inference for Prompt reconstruction), a reference-free and rubric-free reward modeling approach that reformulates reward modeling through backward inference that enables reliable reward modeling in downscaled regimes where judgment methods fail, is proposed.

Yike Wang, Faeze Brahman, Shangbin Feng et al. · 3 citations

Small Updates, Big Doubts: Does Parameter-Efficient Fine-tuning Enhance Hallucination Detection ?

Experimental results show that PEFT consistently strengthens hallucination detection ability, substantially improving AUROC across a wide range of hallucination detectors, and indicates that PEFT methods primarily reshapes how uncertainty is encoded and surfaced, comparing with injecting new factual knowledge into the models.

Xuehai Hu, Yifan Zhang, Song-Tao Wei et al. · 2 citations
#computer vision Preprint Feb 2026

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

UReason is introduced, a benchmark for evaluating reasoning-to-generation alignment in this paradigm, consisting of 2,000 human-curated and human-verified instances spanning five reasoning-intensive tasks: Code, Arithmetic, Spatial, Attribute, and Text, and it is found that decontextualized generation consistently outperforms reasoning-guided generation by a large margin.

Cheng Yang, Chufan Shi, Bo Shui et al. · 5 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.