Experiments show that EASD consistently improves accuracy over standard SD and reward-guided variants while maintaining comparable inference efficiency, suggesting that speculative decoding can serve not only as an acceleration method but also as an effective mechanism for improving reasoning quality.
Tiancheng Su, Meicong Zhang, Guoxiu He· 3 citations
We propose LLM-PeerReview, an unsupervised LLM Ensemble method that selects the most ideal response from multiple LLM-generated candidates for each query, harnessing the collective wisdom of multiple models with diverse strengths. LLM-PeerReview is built on a novel, peer-review-inspired framework that offers a transparent and interpretable mechanism, while remaining fully unsupervised for flexible adaptability and generalization. Specifically, it operates in three stages: For scoring, we use the emerging LLM-as-a-Judge technique to evaluate each response by reusing multiple LLMs at hand; For reasoning, we can apply a straightforward averaging strategy or a principled graphical model-based truth inference algorithm to aggregate multiple scores to produce a final score for each response; Finally, the highest-scoring response is selected as the best ensemble output. LLM-PeerReview is conceptually simple and empirically powerful. Our results across four datasets show that the two variants of the proposed approach outperform the advanced model Smoothie-Global by 6.9% and 7.3% points, cross diverse task types including factual recall QA, math reasoning, and instruction following. Notably, we also establish a carefully curated benchmark suite for LLM Ensemble, integrating 12 methods across four classic datasets and three task families, all evaluated under a rigorous and consistent protocol. We hope this repository will help researchers reproduce the LLM Ensemble baselines.
Zhijun Chen, Zeyu Ji, Qianren Mao et al.· arXiv.org· 5 citations
It is concluded that a human-in-the-loop (HITL) approach is crucial and the production-grade, cloud-based infrastructure designed to support this workflow is presented.
Alessandro M. Lucca, Francesco Corso, Francesco Pierri· arXiv.org· 2 citations
Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance low latency with high translation quality. Despite rapid progress, evaluation remains fragmented across existing frameworks, which make different assumptions about how systems operate - for example, whether they process continuous speech or short pre-segmented audio, and whether they support output revision (retranslation) or not (incremental). For instance, SimulEval, the most widely used framework, supports only incremental decoding, assumes short segmented inputs, and lacks a native support for system demonstrations. As a result, comparing systems fairly and consistently across studies remains challenging, with no unified solution for benchmarking and interactive demonstration. To address this gap, we introduce simulstream, the first open-source framework for StreamST evaluation and demonstration. It supports both incremental and re-translation decoding on long-form speech, provides fine-grained logging for quality and latency evaluation, and includes an interactive web interface for real-time visualization and comparison.
Marco Gaido, Sara Papi, Mauro Cettolo et al.· arXiv.org· 7 citations· ⚡2
Reach audiences
Advertise in front of researchers, engineers, and readers.
SGM is extensible, and its combined defenses, denoted as SGM*, integrate with existing detoxification methods for stronger safety performance, providing an interpretable, low-cost solution for toxicity-controlled multimodal generation.
This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tuned for a specific classification task, and demonstrates its effectiveness on Argumentative Fallacy Detection and Classification tasks, and affective computing tasks on a widely-used dataset.
Nicolas Calbucura, Valentin Barrière· arXiv.org· 1 citation
This work collects data from four different languages, and analyzes how different communicative needs and variations in the listener model influence the informativeness--complexity trade-off, showing that trade-off optimality is not only theoretically achievable but also emerges empirically in learned communication systems.
Phong-Hao Le, Mees Lindeman, Raquel G. Alhama· 0 citations
Overall, internal web-based retrieval functions effectively as a low-latency verification mechanism, but falls short as a reliable IR pipeline, highlighting the need for improved retrieval triggering, query formulation, and evidence-aware confidence calibration in web-enabled LLMs.
PEPPER (PErcePtion-Guided PERturbation), a backdoor defense that rewrites the caption into a semantically distant yet visually similar caption while adding unobtrusive elements, achieves enhanced robustness without training or access to model weights.
Oscar Chew, Po-Yi Lu, Jayden Lin et al.· arXiv.org· 0 citations
A forensic analysis of batch speculative decoding is conducted and it is found that several widely-used implementations silently produce corrupted outputs while reporting competitive speed; failures invisible to metrics like ROUGE.
R. Zhang, Soumik Dey, Ashirbad Mishra et al.· 3 citations
This work introduces a type-induced commitment-communication mechanism that grounds Bayesian Persuasion in natural language dialogue without pre-commitment, and implements two variants: Semi-Formal-Natural-Language (SFNL) and Fully-Natural-Language (FNL), evaluating them against strong baselines and human judges.
Buwei He, Yang Liu, Zhaowei Zhang et al.· arXiv.org· 1 citation
Inspired by recent breakthroughs in Large Language Models (LLMs), this work introduces LLP, the first LLM-based generative framework for second-hand product pricing that substantially surpasses existing methods while generalizing well to unseen categories.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.