TurkmenFST: A Comprehensive Rule-Based Morphological Analysis and Generation System for the Turkmen Language
2026 IEEE Ural-Siberian Conference on Biomedical Engineering, Radioelectronics and Information Technology, USBEREIT 2026, Yekaterinburg, Rusya, 14 - 15 Mayıs 2026, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Doi Numarası: 10.1109/usbereit70063.2026.11580808
- Basıldığı Şehir: Yekaterinburg
- Basıldığı Ülke: Rusya
- Anahtar Kelimeler: low-resource languages, morphological analysis, rulebased system, Turkmen language
- Ankara Üniversitesi Adresli: Evet
Özet
Turkic languages present unique challenges in natural language processing due to their agglutinative morphology, where a single root can theoretically generate thousands of word forms through productive suffixation. While relatively high-resource Turkic languages such as Turkish have well-established morphological tools, the Turkmen language remains critically under-resourced, with no large-scale opensource morphological analyzer available. This paper presents a comprehensive, open-source, rule-based morphological analysis and generation system for Turkmen. The system is built on a modular Python architecture comprising five core components: a phonological rule engine, a morphotactic finite-state machine, a generation module, a generator-verified analysis module, and a lexicon manager. A lexicon of 32,738 entries was compiled from five independent sources through a nine-stage validation pipeline, constituting the largest open-source morphologically tagged stem dictionary reported for Turkmen. The morphological engine models 19 inflectional codes covering seven conjugated tenses, five moods, four participle and converb forms, and three voice derivations, together with composable phonological rules for vowel harmony, consonant softening, vowel elision, and labial rounding. Corpus-based evaluation on a 658,881-token news corpus from the official Turkmen news agency yielded 96.37 percent token coverage and 70.84 percent type coverage, substantially exceeding previously reported figures. External cross-validation against an authoritative national language portal confirmed complete headword coverage across 20,120 entries and identified eight verb paradigm errors that were subsequently corrected. The system and its lexicon are publicly available as open-source resources.