Modeling Surface-Complexity-Based Readability Levels in Turkish via Part-of-Speech Profiles
This study investigates the relationship between part-of-speech (POS) profiles and surface-complexity-based readability levels in Turkish texts. These levels are derived solely from a surface-based proxy complexity score rather than from expert human readability judgments. Twenty-seven POS-based features were extracted from a corpus of 19,779 sentences annotated with 14 POS tags by three annotators with high inter-annotator agreement (Fleiss kappa: 0.84). Kruskal-Wallis tests showed that 25 of 27 features significantly differ across readability levels. The most discriminative features were noun-verb ratio, POS bigram diversity, and noun ratio. A POS-only model achieved 35 percent macro-F1 in five-class classification, well above random baseline. Results demonstrate that Turkish sentence-level complexity correlates meaningfully with POS composition and sequential structure beyond mere length measures.