A Multidimensional Evaluation of Al-Powered Medical Translation Systems: BLEU-AHP Model and User Behavior Insights
This study evaluates six AI medical translation systems using a mixed-methods approach, integrating BLEU scores, user surveys (N=775), and behavioral data. A standardized bilingual corpus was constructed from authoritative sources including the WHO and NMPA, while an AHP-BLEU hybrid model was developed to combine subjective user evaluations with objective scores across word, sentence, and paragraph levels in both Chinese-English and English-Chinese tasks. Results show Atman and Youdao outperform others in overall quality, with DeepL excelling in terminology. Spearman correlation analysis confirms a strong positive association (ρ=0.943, p=0.005) between BLEU scores and user satisfaction, validating the model. Despite rapid advances, current AI medical translation tools still struggle with term accuracy, context adaptation, and document complexity. The proposed AHP-BLEU framework helps align evaluation with user priorities, offering a more balanced view of performance. Future improvements should include semantic-aware metrics and human-verified baselines to better support multilingual medicine applications, from Traditional Chinese Medicine globalization to virtual consultations.