Evaluation of automatic simplification of German-language literary texts using large language models
Abstract
The aim of this study is to investigate the applicability of large language models for the simplification of German-language literary texts. The paper provides a brief overview of existing approaches to this task, as well as statistical metrics used to assess text complexity and readability. Two types of statistical analysis (analysis of variance and correlation analysis) were applied to the values of a range of quantitative indicators obtained from original and simplified texts. These methods were used to test the hypothesis concerning the existence of differences between different groups of texts, as well as the relationships between the analyzed parameters. The scientific novelty of the study lies in the first comprehensive evaluation of the generated simplified texts based on thirteen quantitative indicators. The study also demonstrates that, in the process of simplifying German-language literary texts with large language models, changes in lexical density do not exhibit statistically significant differences. The results of the analysis show that texts simplified using the Llama 3 and MarianMT models are characterized by vocabulary corresponding to lower levels of the CEFR scale. Furthermore, the Llama 3 model significantly reduces text length during the simplification process.