Skip to content
Open access

The Role of Lexical and Semantic Features in Shaping Indonesian Academic Text Readability: An NLP-Based Analysis

Jul 2026 · Jurnal Sosioteknologi · Vol 25 · 0 citations · 39 references

TL;DR

The results show that semantic features consistently exhibit stronger relationships with readability than conventional lexical measures, indicating that semantic clarity has a greater impact on the readability of Indonesian academic writing than lexical sophistication alone.

Abstract

This study investigates the relationships among lexical features, semantic features, and readability using a corpus of 27,274 sentences extracted from Indonesian undergraduate academic abstracts. A quantitative descriptive-correlational design was employed, combining natural language processing (NLP)-based feature extraction with a regression-calibrated Indonesian adaptation of the Flesch readability formula. Pearson and Spearman correlation analyses were performed to examine associations among lexical density, lexical diversity, lexical bundles, semantic dimensions, and readability scores. The results show that semantic features consistently exhibit stronger relationships with readability than conventional lexical measures. Mean word concreteness exhibited the strongest positive correlation with readability (r = 0.424), whereas the abstractword ratio showed the strongest negative correlation (r = -0.407). In contrast, lexical density, lexical diversity, and lexical bundles displayed only negligible to weak correlations. These findings indicate that semantic clarity has a greater impact on the readability of Indonesian academic writing than lexical sophistication alone. Rather than emphasizing increasingly complex vocabulary, improving readability depends on communicating ideas through concrete and contextually meaningful language. Beyond extending Indonesian readability research, this study provides empirical evidence for the development of NLP-based readability assessment, intelligent academic writing assistance systems, and pedagogical practices that support clearer scholarly communication for national and international audiences.

Read PDF

Similar papers

Preprint Aug 2026

Machine learning and digital pragmatics: Which word category influences emoji use most?

This study examines the performance of the state-of-the-art MARBERT model in identifying the lexical/pragmatic category associated with emoji use on X within a digital pragmatics approach (DPA). A net corpus of 15856 Colloquial Arabic (CA) posts containing emojis was collected from X using Python. The texts were tokenized and normalized into 4 lexical categories, namely noun_norm, verb_norm, adj_norm, and adverb_norm, and 2 pragmatic/structural categories, question_norm and exclamation_norm. MARBERT was finetuned and optimized to identify which category scores standard metrics more, hence associated with emoji use, while binary logistic regression was used to examine which category is statistically associated with emoji occurrence. Findings unveil that nouns dominate the corpus in normalized frequency (M = 0.675, SD = 0.161), followed by verbs (M = 0.083, SD = 0.100). However, verbs have the strongest influence of emoji use indicated by verb density (\b{eta} = 0.821, p = .001, 95% CI [0.332, 1.309]). The study concludes that in digital pragmatics of CA on X, emoji use association with lexical/pragmatic category can be explained by a hybrid approach of computational, statistical, and pragmatic methods, reflecting the interaction among machine learning, linguistic/lexical features, contextual representation, and pragmatic communication.

Mohammed Q. Shormani, Y. A. AlSohbani · 0 citations
Open access 2026

Characterization and Mechanisms of Lexical Complexity in AI-Generated Texts: A Comparative Corpus-Based Study

: Based on a corpus-based methodology, this study analyzes the intrinsic reasons for the high level of lexical complexity observed in Artificial Intelligence Generated Content (AIGC). The research compares 24 English argumentative essays written by AI with 24 second-language (L2) learner essays reaching the IELTS Writing Task 2 Band 7 level. Under controlled conditions of identical genre and topic, the study performs quantitative statistics across three dimensions: lexical sophistication, semantic abstraction, and information density. Statistical results indicate that the frequency of advanced vocabulary in AI texts is significantly higher, approximately 2.3 times that of human texts. The proportion of abstract nouns reached 9.14%, far exceeding the 3.32% found in human texts, suggesting that AI expressions tend toward nominalization and conceptualization. Regarding overall information organization, the lexical density of AI texts was 69.9%, also surpassing the 60.3% of human texts, reflecting a stronger tendency for information condensation and phrasal structures. The analysis points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions. This mechanism favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface. Such complexity is essentially a formal feature at the statistical level and is not entirely equivalent to the proficiency levels corresponding to human L2 acquisition. These findings provide empirical references for AI text identification, the refinement of writing evaluation standards, and L2 writing pedagogy.

Lulu Chen · 0 citations
Open access Jul 2026

Analysis of Morphological and Semantic Transformations of English Words with Indonesian Affixes in Digital Communication

Nowadays, the increasing use of English words combined with Indonesian affixes in digital communication has given rise to a unique form of linguistic innovation that reflects the interaction between the local and global language. This research aims to analyze how Indonesian affixes influence the morphological structure and semantic meaning of English words which are used in digital communication. This research uses qualitative descriptive to analyze 45 lexical items found from Instagram and YouTube comment sections. The data selected from real online interactions containing English words attached to Indonesian affixes. Data were analyzed by following four stages, such as data reduction, classification, interpretation, and conclusion. The findings show that the Indonesian suffix -nya is the most frequently used representing 66.67% of the collected data, followed by the prefixes di-, ke-, nge-, ter-, se-, and the suffix -an. These affixes modify English words by adapting them to Indonesian grammatical structures while preserving or extending their meanings. Unlike previous studies that primarily examined English-Indonesian lexical borrowing through corpus data or magazine texts, this study contributes new insights by examining the morphological and semantic adaptation of English lexical items in spontaneous digital interactions on Instagram and YouTube.  The research determines that Indonesian-English hybrid expressions showcase linguistic innovation, facilitate efficient digital interaction, and exemplify the active development of bilingual language usage in modern Indonesian online conversations.

Prihatin Puji Astuti, Ria Antika · 0 citations
Open access Aug 2026

Semantic analysis of problems in natural language processing and their mathematical interpretation

Semantic analysis has become a central challenge in natural language processing, driven by exponential growth in digitized textual data and the need for automated content processing across multiple applications including machine translation, text classification, sentiment analysis, and information retrieval. However, while semantic analysis methods are well-developed for resource-rich languages such as English, morphologically complex languages like Uzbek suffer from deficiencies in annotated corpora, lexical-semantic resources, and high-quality vector models – a gap amplified by governmental initiatives in digital economy development and national language technology advancement. This section grounds semantic analysis in the distributional semantics hypothesis principle that words exhibiting similar contexts possess similar meanings – thereby recasting the problem as a geometric challenge within continuous vector spaces. Two principal mathematical strategies are formalized: (1) prediction-based models (word2vec: CBOW/Skip-gram), which optimize context prediction objectives, and (2) count-based models (GloVe), which leverage global co-occurrence statistics through matrix factorization. Both project high-dimensional word co-occurrence relationships into low-dimensional dense vector spaces, enabling semantic analogy representation. For resource-scarce languages like Uzbek, cross-lingual embedding alignment (Procrustes optimization) enables semantic knowledge transfer from resource-rich languages, facilitating shared semantic spaces across the Turkic language family. The section concludes with formal problem specification: given vocabulary V and corpus C, semantic analysis is formalized as (1) a mapping problem preserving distributional properties, (2) an optimization problem minimizing loss through gradient-based methods, and (3) an evaluation problem assessing quality through semantic similarity, analogy, and downstream NLP task performance.

D. Akhmedjanova · 0 citations
Open access Jul 2026

Lexical diversity and CEFR vocabulary in critical academic reading passages: A conceptual replication

Background and Purpose: Critical academic reading skills demand learners to engage in higher-order reading skills which require advanced lexical knowledge. The two dimensions of lexical knowledge, namely lexical diversity and vocabulary variety, however, are mostly assessed in isolation using unstable measurement instruments, hence giving misleading results. The purpose of this study was to examine these two dimensions in critical academic reading passages through an integrated approach using a more reliable measurement instrument. Methodology: The corpus consisted of 32 selected reading passages. They were first cleaned to remove any interference, such as numbering that might affect the lexical calculation, before being converted to plain text format. The Text Inspector was then used to generate the Measure of Textual Lexical Diversity (MTLD) and CEFR vocabulary difficulty levels for all passages. The scores were then analysed descriptively to examine the patterns of lexical diversity and CEFR vocabulary distributions across the passages. Additionally, Spearman’s rho correlational analysis was conducted to examine the relationship between MTLD scores and CEFR-based vocabulary levels. Findings: All 32 examined passages generally exhibit high lexical diversity, with many common and basic words. Lexical diversity was found to increase when advanced vocabulary was added to the passage, indicating a significant positive correlation between the two dimensions. Contributions: The integrated approach in lexical analysis adopted in this study results in a comprehensive picture of lexical complexity rather than using the single metric alone, hence offering a clearer and fairer approach for instructors in evaluating and selecting reading passages to be used with students. Keywords: CEFR profiling, critical academic reading, lexical difficulty, lexical diversity, Measure of Textual Lexical Diversity (MTLD). Cite as: Aziz, A., Syed Ahmad, T. S. A., Awang, S., Abdul Aziz, R., Azlan, N. A., & Ahmad, S. N. (2026). Lexical diversity and CEFR vocabulary in critical academic reading passages: A conceptual replication. Journal of Nusantara Studies, 11(2), 228-242. https://dx.doi.org/10.24200/jonus.vol11iss2pp228-242

Anealka Aziz, Tuan Sarifah Aini Syed Ahmad, Suryani Awang et al. · 0 citations
Conference Jul 2026

Modeling Surface-Complexity-Based Readability Levels in Turkish via Part-of-Speech Profiles

This study investigates the relationship between part-of-speech (POS) profiles and surface-complexity-based readability levels in Turkish texts. These levels are derived solely from a surface-based proxy complexity score rather than from expert human readability judgments. Twenty-seven POS-based features were extracted from a corpus of 19,779 sentences annotated with 14 POS tags by three annotators with high inter-annotator agreement (Fleiss kappa: 0.84). Kruskal-Wallis tests showed that 25 of 27 features significantly differ across readability levels. The most discriminative features were noun-verb ratio, POS bigram diversity, and noun ratio. A POS-only model achieved 35 percent macro-F1 in five-class classification, well above random baseline. Results demonstrate that Turkish sentence-level complexity correlates meaningfully with POS composition and sequential structure beyond mere length measures.

Zühal Calayoğlu, Özkan Aslan · 0 citations