Dergiler / Turkish Journal of Electrical Engineering and Computer Sciences / 2019 / Cilt: 27 - Sayı: 2

Turkish lexicon expansion by using finite state automata

Sayfa
1012–1027
DOI
—

Abstract

Turkish is an agglutinative language with rich morphology. A Turkish verb can have thousands of differentword forms. Therefore, sparsity becomes an issue in many Turkish natural language processing (NLP) applications.This article presents a model for Turkish lexicon expansion. We aimed to expand the lexicon by using a morphologicalsegmentation system by reversing the segmentation task into a generation task. Our model uses finite-state automata(FSA) to incorporate orthographic features and morphotactic rules. We extracted orthographic features by capturingphonological operations that are applied to words whenever a suffix is added. Each FSA state corresponds to either astem or a suffix category. Stems are clustered based on their parts-of-speech (i.e. noun, verb, or adjective) and suffixesare clustered based on their allomorphic features. We generated approximately 1 million word forms by using only a fewthousand Turkish stems with an accuracy of 82.36%, which will help to reduce the out-of-vocabulary size in other NLPapplications. Although our experiments are performed on Turkish language, the same model is also applicable to otheragglutinative languages such as Hungarian and Finnish.