Skip to content
Open access

The Lexical Analysis of Postgraduate Artificial Intelligence Academic Texts

Abstract

This thesis investigates the vocabulary and incidental vocabulary learning opportunities in authentic academic texts for postgraduate students of Artificial Intelligence. Book chapters and journal articles from five courses for taught masters of Artificial Intelligence at Victoria University of Wellington were collected to compile the corpus of Artificial Intelligence reading texts (CAIRT). The thesis consists of three studies. The first study investigates the vocabulary profile of texts in CAIRT, assessing how much vocabulary is needed to reach 95% and 98% coverage using Nation’s (2012) 25 BNC/COCA word lists with five supplementary lists. The findings show that 4,000 and 6,000 word families plus supplementary lists are needed to reach 95% and 98% coverage, respectively. However, lexical demands vary across courses and text types, indicating that the vocabulary load cannot be generalised even within a single academic discipline. Additionally, the first 3,000 word families account for the largest proportion of the corpus, while mid-frequency word families and supplementary-list words make comparable contributions, with low-frequency word families accounting for the smallest proportion. The second study explores the repetition and distribution of word families from each category (high-, mid-, low-frequency, and supplementary lists). Although high-frequency word families are most likely to recur, the majority of word families occur only a small number of times. While many word families are shared across courses and trimesters, a substantial proportion are restricted to individual sub-corpora, indicating uneven opportunities for incidental vocabulary learning across the texts and courses. The third study investigates which words are elaborated within texts that may facilitate vocabulary learning and reading comprehension. Using Hyland’s (2005) taxonomy of code glosses, 57 single words and 188 multiword units (MWUs) are identified as elaborated within two courses. The elaborated single words include high-, mid-, and low-frequency words, abbreviations, proper nouns, transparent compounds, and other words outside the BNC COCA 25,000 word families. The corpus frequency rate of these elaborated words varies, and their distribution is inconsistent across texts. Many of these elaborated words occur in only one text, and a small number of them recur across multiple texts. Five single words and seven of the MWUs are elaborated multiple times, and most of the elaboration occurs within a single text. Overall, the findings demonstrate that the lexical demands of authentic postgraduate AI readings are more nuanced than estimates of vocabulary coverage alone imply and cannot be generalised even within a single discipline. Although authentic disciplinary texts repeatedly recycle a relatively small core vocabulary, they provide uneven opportunities for incidental vocabulary learning through repetition and lexical elaboration. Compared with the adapted reading materials and authentic texts used in experimental studies of incidental vocabulary learning through reading, authentic AI readings provide these lexical conditions less consistently, as their primary purpose is to communicate disciplinary knowledge rather than to facilitate vocabulary learning. These findings contribute to a more contextualised understanding of incidental vocabulary learning in authentic disciplinary reading and have implications for vocabulary profiling, course sequencing, and pedagogical support for postgraduate AI students.

Read PDF