Skip to content
Open access

Nominalization Patterns in Tombatu Language: A Generative Morphological Analysis of Affixal Noun Formation

Jul 2026 · REiLA: Journal of Research and Innovation in Language · 0 citations

Abstract

Regional languages preserve complex grammatical systems that reveal how communities organize experience, identity, and cultural knowledge. However, nominalization in Tombatu, an underdescribed Austronesian language of Minahasa, has not been systematically examined through an explicit rule-based morphological framework. This study addresses this gap by investigating the structural patterns, semantic functions, and generative mechanisms of affixal noun formation in Tombatu. Employing a descriptive-taxonomic design, data were collected over six months in Tombatu Village from proficient speakers through naturally occurring speech, elicitation, recording, note-taking, and informant validation. The final dataset found unique nominalized forms. These forms were classified into three major processes comprising 11 noun-forming prefixes, one noun-forming suffix, and 14 affix combinations, yielding 26 identified affixal configurations. The analysis applied Word Combination Rules, Derivation Rules, the Boundary Insertion Convention, and relevant morphological constraints. The findings show that prefixation exhibits the greatest structural and semantic variation in the dataset, producing agentive, person-denoting, instrumental, object-related, and result-related nouns. For example, /tapə-/ combines with /lukuʔ/ ‘to drink’ to form /tapəlukuʔ/ [tapəlukuʔ] ‘drinker’. Suffixation is more restricted, with /-an/ forming locative nouns, as in /tawoi/ ‘to work’ plus /-an/, which produces /tawoian/ [tawoian] ‘workplace’. Complex affix combinations form nouns denoting places, processes, objects, states, results, and entities. This study provides the first systematic generative account of Tombatu nominalization, expands empirical evidence for Austronesian word formation, strengthens regional-language documentation, and offers a linguistic basis for language maintenance and potential contrastive morphological activities in multilingual education.

Read PDF

Similar papers

Open access Aug 2026

Major Word Classes in the Lohorung Language: A Functional Typological Sketch

This study examines the four major word classes (viz. noun, verb, adjective and adverb) in Lohorung, a Kirat Rai (Tibeto-Burman) language of north-eastern Nepal, based on their semantic, syntactic, and morphological properties, based on Givón’s formal and functional approach. Data were collected through elicitation with native speakers and textual analyses and were analysed with reference to the author’s native-speaker intuition. The findings indicate that Lohorung exhibits prototypical nouns with their multiple features and non-prototypical nouns as well. Morphologically, nouns are marked by the specific affixes, <-ʈsi> encodes dual and <-i> encodes plural. Lohorung employs three types of classifiers, namely <-tsi> <-kɔ>, and <-pɑŋ/pɑ>. Verbs are clause-final and serve as main predicates, typical of Tibeto-Burman languages. Lohorung has three numbers and persons systems with clusivity. Manner adverbials may take , and some adverbs modify adjectives. This study concludes that the analysis of the four major word classes enhances the grammatical description of the Lohorung language and contributes to functional-typological research on lexical classification.

Diwas Rai · 0 citations
Open access Jul 2026

FORMAL AND NEURAL APPROACHES TO MORPHOLOGICAL ANALYSIS IN AGGLUTINATIVE LANGUAGES: EVIDENCE FROM KAZAKH

Automatic morphological analysis remains a challenging task for agglutinative languages because of their rich inflectional systems, productive derivation, and complex morphophonological rules. Recent neural models, especially transformer-based architectures, have demonstrated impressive empirical performance. However, they frequently exhibit a deficiency in linguistic transparency and encounter challenges in systematic generalization inside low-resource environments. This research offers a comparative and integrative examination of formal (rule-based and finite-state) and neural (KazBERT-based) methodologies for morphological analysis, utilizing the Kazakh language as a case study of low-resource agglutinative morphology. Initially present a formal morphological model that distinctly represents root-affix structure, vowel harmony, and morphotactic restrictions. We next test many neural architectures for morphological disambiguation and tagging, such as KazBERT coupled with CRF-based decoding. In addition to typical accuracy measurements, we do a comprehensive error taxonomy and linguistic analysis, investigating how various model classes manage ambiguity, infrequent forms, and extended affix chains. The findings indicate that whereas neural models excel in surface-level accuracy compared to exclusively rule-based systems, they demonstrate consistent deficiencies in morphologically intricate and infrequent constructs. On the other hand, formal models show better generalization based on language limitations. Based on these results, we suggest a hybrid morphology-aware framework that adds symbolic restrictions to neural inference. This framework consistently improves results in a variety of assessment contexts. The study demonstrates that effective morphological analysis of agglutinative languages requires the integration of neural representation learning with explicit linguistic structure. The results are not tied to any one language and have wider implications for morphology-sensitive NLP in low-resource settings.

A. Aitim, Ә.Қ. Әйтім, Халықаралық Ақпараттық et al. · 0 citations
Review Open access Jul 2026

Linguistic Typology: A Critical Review of Metacoding and Syntactic Structure across Languages

Linguistic typology seeks to describe the structural diversity of the world's languages and to explain the recurrent patterns, or universals, that constrain this diversity. A central but often under-examined concept within this enterprise is metacoding: the principled allocation of formal marking resources, such as case, agreement, and word order, across grammatical relations and semantic or pragmatic categories according to markedness, economy, and iconic motivation. This review synthesises recent typological scholarship on metacoding and its bearing on syntactic structure across languages, drawing on comparative concepts, differential argument marking, alignment typology, morphological complexity trade-offs, and the growing role of large-scale computational databases in typological inference. The review traces how metacoding principles account for asymmetries in the marking of subjects and objects, for the distribution of ergative and accusative alignment, and for the coevolution of nominal and verbal grammatical marking. It further considers how phylogenetic and areal methods have reshaped debates about whether word-order correlations reflect cognitive universals or lineage-specific historical accidents, and how the loss of linguistic diversity through language endangerment threatens the empirical base on which such debates depend. Methodological issues concerning sampling, representativeness, and the comparability of descriptive categories across languages receive particular attention, since these issues bear directly on the reliability of typological generalisations. The review concludes that metacoding offers a unifying, functionally motivated account of why languages allocate overt marking unevenly across categories, while cautioning that the evidential basis for many typological universals remains uneven across language families and geographical regions. Future research directions, together with the principal limitations of the current evidence base, are outlined.

Sun Jiaze, Mulyadi, Khairina et al. · 0 citations
Preprint Aug 2026

Rethinking and formalising the state across languages: a unified computational learning theory account

It is suggested that the state is a systemic, context-dependent morphosyntactic mechanism that selects grammatical templates across synthetic languages and constitutes one instance of a broader class of syntactically conditioned dependencies that also includes agreement and grammatical case.

M. E. Idrissi · 0 citations
Open access Jul 2026

The morphological origins and semantic diversity of English compound verbs

English compound verbs are said to result predominantly from a morphological operation of conversion or desuffixation or, more rarely and ‘directly’, from an operation of compounding. In order to confirm the existence of ‘direct’ compound verbs in contemporary English and assess their relative weight among all compound verbs, a sample set of 250 units was obtained by automatically extracting recent verbal units containing a hyphen from the Oxford English Dictionary. Their dates of first attestation were then compared with those of similar nominal units, and those verbal units that were not clearly preceded by nominal occurrences were considered likely ‘direct’ compounds. They appear to constitute a sizable minority of the sample (almost one-fifth of all units). In a second step, in order to gain insight into the variety and quantitative distribution of the semantic relations attested in the sample, the internal semantic relation of each unit was established, applying, where possible, Pepper’s (2023) list of relations originally developed for binominal units. This showed that the verbal nature of the compounds leads to the presence of two additional, specific relations, MANNER and ATTRIBUTION, and that the morphological origin of the compound verb has an influence on the dominant categories of internal relations.

Pierre J. L. Arnaud, Vincent Renner · 0 citations
Preprint Jul 2026

Annotating Korean adnominal ending constructions in corpus data: Beyond relative-clause identification

The Korean adnominal ending \texttt{ETM} occurs in diverse noun-modifying constructions, including relative-clause-like modifiers, adjectival and copular forms, bound-noun constructions, and lexicalized expressions. This paper argues that \texttt{ETM} is not a direct marker of relative-clause structure, but a morphological exponent shared by several adnominal constructions. We propose a corpus-based typology that distinguishes these constructions using predicate type, auxiliary structure, argument-structural compatibility, head-noun restriction, and lexicalized patterns. We operationalize the typology as a construction-sensitive annotation layer for the KLUE dependency treebank, implemented through an ordered rule-based procedure and evaluated by manual validation. Productive relative-clause-like uses account for 39.4\% of the analyzed instances; the remainder consists mainly of adjectival, copular, bound-nominal, modal, temporal, and collocational constructions. The findings show that Korean relative-clause-like modification cannot be identified from adnominal morphology alone.

Jungyeul Park, Chulwoo Park · 0 citations