Skip to content

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

It is shown that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling, are effective for segment-level SI evaluation.

Abstract

Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is designed for rubric-aligned segment-level SI evaluation. We construct a professionally annotated corpus of 1,101 SI segments with scores for meaning transfer (LQ), delivery quality (EXP), and perceived latency (LAT). We show that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling. To isolate supervision structure under identical backbone capacity, we introduce dual regression heads on a LoRA-adapted COMET-KIWI encoder. On a held-out talk-level test set, the model achieves Pearson correlations of 0.388 (LQ) and 0.301 (EXP), improving over frozen COMET-KIWI. Given low absolute rater agreement, we interpret results relative to human consistency and target stable ranking signals for formative assessment.

View source

Similar papers

#natural language process... Preprint Sep 2026

EviSI: An Evidence-Based Evaluation Agent for Simultaneous Interpreting

Low-latency simultaneous speech-to-speech translation must keep pace with ongoing speech while preserving key information. To meet these demands, systems use segmentation, reformulation and condensation to reorganize and rephrase information. However, metrics developed for text translation, including BLEU and COMET, ma...

Ben Yan, Zong-Yao Li, Xiao-Yu Chen et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
Open access Aug 2026

Towards Trustworthy Large Language Models

An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.

Sakshi Parate, Shreyans Sanyal · 0 citations
#machine learning Preprint Aug 2026

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks is proposed, a behavioral diagnostic that separates failure diagnosis from abstention scoring.

Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty · 0 citations
Conference Aug 2026

Evaluating AutoEDA Without Human Raters: A Diagnostic Framework and Case Study of Structured Question Generation

Human evaluation remains the dominant way to assess automated exploratory data analysis (AutoEDA) systems, but it is expensive, subjective, and hard to reproduce. We introduce a reproducible automatic diagnostic framework that surfaces structural and statistical failures in AutoEDA—failure modes often under-emphasized...

B. N. Do, Phu-Vinh Nguyen, Hung-Nghiep Tran · 0 citations
#natural language process... Preprint Sep 2026

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an...

Hritika Sharma, Thibault Bañeras-Roux, Alessandra Pinto et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.