Skip to content

Category

natural language processing

3,089 papers

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

Self-Evaluation Elicitation (SEE) is introduced, a method that surfaces a latent ability to predict how a judge will score its own output through a short cycle comprising a calibration-coupled reinforcement learning phase that improves the answer and predicts the judge, followed by a masked distillation phase that sharpens the prediction while leaving the answer untouched.

XiuYu Zhang, Yingyu Shan, Junfeng Fang et al. · 2 citations

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

RealClawBench is introduced, a live benchmark framework built from real OpenClaw sessions to capture the distribution, diversity, and real-world difficulty of deployed agent use and provides a practical path toward benchmarks that better measure agent capability in actual use.

Zongwei Lv, Zhewen Tan, Yao-Ming Li et al. · 1 citation · ⚡1

CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning

By isolating task-specific patterns into independent modules, CRAM mitigates catastrophic forgetting across tasks and boost parameter efficiency, and utilizes adaptive-rank instantiation to identify the capability gap between existing expert capability and new task demands, and dynamically allocate only the necessary parameters.

Jun Tang, Zhen Xie, Yucheng Shi et al. · 0 citations

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

This work introduces LongJudgeBench, a comprehensive benchmark for evaluating LLM judges on long-form outputs across diverse real-world scenarios and judging protocols, and systematically evaluates a broad range of LLM judges, covering multiple base models and judging settings.

Junjie Chen, Yuxin Dong, Haitao Li et al. · 0 citations

Auditing LLM Benchmarks with Item Response Theory

An Item Response Theory-based indicator is introduced that surfaces likely mislabels at 95% precision in the top 200 examples across seven preference and multiple-choice benchmarks using responses from 114 models, outperforming a supervised classifier.

Sander Land, D. Bikel · 2 citations

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

Overall, the results show that HLV can be learned as annotator-specific label-explanation behavior, suggesting a path toward scalable explanation-based annotation grounded in annotator histories rather than labels alone.

Beiduo Chen, Pingjun Hong, Ziyun Zhang et al. · 0 citations
#natural language process... Conference Open access May 2026

Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?

This work examines whether large language models exhibit similar behaviors when assigned high or low status personas, and shows that LLMs show key socio-cognitive effects of power, albeit with nuances and variability, linking simulated interactions to both desirable and unsafe behaviors.

Anvesh Rao Vijjini, S. Manjunath, Snigdha Chaturvedi · 0 citations

Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects

It is concluded that current transformer models do not explain human morphosyntactic processing, and that evaluations of transformers as cognitive models must adopt rigorous, comprehensive experimental designs to avoid spurious generalizations from isolated syntactic configurations or individual models.

Titus von der Malsburg, Sebastian Padó · 2 citations

The Company You Keep: How LLMs Respond to Dark Triad Traits

This work examines how LLMs respond to user prompts expressing varying degrees of Dark Triad traits (Machiavellianism, Narcissism, and Psychopathy) using a curated dataset, revealing systematic differences across models.

Zeyi Lu, A. Henestrosa, Pavel Chizhov et al. · 1 citation

Beyond the Rabbit Hole: Mapping the Relational Harms of QAnon Radicalization

Large-scale computational research on conspiracy theories has focused exclusively on believers'online behavior, leaving the harm experienced by those closest to them under-examined. This paper bridges this gap by analyzing 12747 stories from r/QAnonCasualties, an online support group for people who have ``lost''someone to conspiracy beliefs. We design a computational pipeline to extract fine-grained thematic traits from personal narratives and cluster them into six coherent radicalization personas, which we then link to the emotional toll reported by narrators via LLM-assisted emotion detection and regression modeling. We find that personas are meaningful predictors of specific emotional harms: radicalization perceived as a deliberate ideological choice is associated with anger and disgust, while personas marked by personal and cognitive collapse correspond to fear and sadness. This work provides an empirically grounded computational framework for understanding the relational harms of radicalization, opening new avenues for research into its wider social consequences.

Ngoc Bich Doan, Giuseppe Russo, Gianmarco De Francisci Morales et al. · 0 citations

Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations

We explore intrinsic dimension (ID) of LLM representations as a marker of linguistic complexity. Specifically, we test whether ID differences across model layers reflect well-known complexity contrasts established in (psycho)linguistics: coordination vs. subordination, right-branching vs. center-embedding, and unambiguous vs. ambiguous attachment. Our results on six different LLMs show that these contrasts are consistently reflected in ID differences, with more complex phenomena eliciting higher ID profiles. Notably, ID differences emerge at different points across layers for different contrasts, also reaching their peaks at different stages. Further experiments using representational similarity and layer pruning confirm the trends. We conclude that ID is a useful marker of linguistic complexity in LLMs, that it points to similar linguistic processing steps across disparate LLMs, and that it has the potential to differentiate between different types of complexity.

Marco Baroni, Emily Cheng, Iria deDios-Flores et al. · 3 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.