Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting
It is shown that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling, are effective for segment-level SI evaluation.
Zi-Yu Zhang, Satoshi Nakamura
· 0 citations