Skip to content

Category

natural language processing

2,394 papers

Pearmut: Human Evaluation of Translation Made Trivial

Pearmut is introduced, a lightweight yet feature-rich platform that makes end-to-end human evaluation as easy to run as automatic evaluation and enables reliable human evaluation to become a practical, routine component of model development and diagnosis rather than an occasional effort.

Vilém Zouhar, Tom Kocmi · 11 citations · ⚡2
#artificial intelligence Preprint Dec 2025

Entropy-Aware Token Rejection for Improving Speculative Decoding

Experiments show that EASD consistently improves accuracy over standard SD and reward-guided variants while maintaining comparable inference efficiency, suggesting that speculative decoding can serve not only as an acceleration method but also as an effective mechanism for improving reasoning quality.

Tiancheng Su, Meicong Zhang, Guoxiu He · 3 citations
#artificial intelligence Review Dec 2025

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process

We propose LLM-PeerReview, an unsupervised LLM Ensemble method that selects the most ideal response from multiple LLM-generated candidates for each query, harnessing the collective wisdom of multiple models with diverse strengths. LLM-PeerReview is built on a novel, peer-review-inspired framework that offers a transparent and interpretable mechanism, while remaining fully unsupervised for flexible adaptability and generalization. Specifically, it operates in three stages: For scoring, we use the emerging LLM-as-a-Judge technique to evaluate each response by reusing multiple LLMs at hand; For reasoning, we can apply a straightforward averaging strategy or a principled graphical model-based truth inference algorithm to aggregate multiple scores to produce a final score for each response; Finally, the highest-scoring response is selected as the best ensemble output. LLM-PeerReview is conceptually simple and empirically powerful. Our results across four datasets show that the two variants of the proposed approach outperform the advanced model Smoothie-Global by 6.9% and 7.3% points, cross diverse task types including factual recall QA, math reasoning, and instruction following. Notably, we also establish a carefully curated benchmark suite for LLM Ensemble, integrating 12 methods across four classic datasets and three task families, all evaluated under a rigorous and consistent protocol. We hope this repository will help researchers reproduce the LLM Ensemble baselines.

Zhijun Chen, Zeyu Ji, Qianren Mao et al. · 5 citations

Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems

Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance low latency with high translation quality. Despite rapid progress, evaluation remains fragmented across existing frameworks, which make different assumptions about how systems operate - for example, whether they process continuous speech or short pre-segmented audio, and whether they support output revision (retranslation) or not (incremental). For instance, SimulEval, the most widely used framework, supports only incremental decoding, assumes short segmented inputs, and lacks a native support for system demonstrations. As a result, comparing systems fairly and consistently across studies remains challenging, with no unified solution for benchmarking and interactive demonstration. To address this gap, we introduce simulstream, the first open-source framework for StreamST evaluation and demonstration. It supports both incremental and re-translation decoding on long-form speech, provides fine-grained logging for quality and latency evaluation, and includes an interactive web interface for real-time visualization and comparison.

Marco Gaido, Sara Papi, Mauro Cettolo et al. · 7 citations · ⚡2

SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification

SGM is extensible, and its combined defenses, denoted as SGM*, integrate with existing detoxification methods for stronger safety performance, providing an interpretable, low-cost solution for toxicity-controlled multimodal generation.

Hongbo Wang, Maungmaung Aprilpyone, Isao Echizen · 1 citation

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification

This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tuned for a specific classification task, and demonstrates its effectiveness on Argumentative Fallacy Detection and Classification tasks, and affective computing tasks on a widely-used dataset.

Nicolas Calbucura, Valentin Barrière · 1 citation
#artificial intelligence Preprint Nov 2025

On the Optimality of Kinship Naming: an Information-theoretic Approach

This work collects data from four different languages, and analyzes how different communicative needs and variations in the listener model influence the informativeness--complexity trade-off, showing that trade-off optimality is not only theoretically achievable but also emerges empirically in learned communication systems.

Phong-Hao Le, Mees Lindeman, Raquel G. Alhama · 0 citations

Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

Overall, internal web-based retrieval functions effectively as a low-latency verification mechanism, but falls short as a reliable IR pipeline, highlighting the need for improved retrieval triggering, query formulation, and evidence-aware confidence calibration in web-enabled LLMs.

Sahil Kale · 1 citation

PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models

PEPPER (PErcePtion-Guided PERturbation), a backdoor defense that rewrites the caption into a semantically distant yet visually similar caption while adding unobtrusive elements, achieves enhanced robustness without training or access to model weights.

Oscar Chew, Po-Yi Lu, Jayden Lin et al. · 0 citations

Make an Offer They Can't Refuse: Grounding Bayesian Persuasion in Real-World Dialogues without Pre-Commitment

This work introduces a type-induced commitment-communication mechanism that grounds Bayesian Persuasion in natural language dialogue without pre-commitment, and implements two variants: Semi-Formal-Natural-Language (SFNL) and Fully-Natural-Language (FNL), evaluating them against strong baselines and human judges.

Buwei He, Yang Liu, Zhaowei Zhang et al. · 1 citation

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.