Journals / Gümüşhane Üniversitesi Fen Bilimleri Dergisi / 2018 / Cilt: 8 - Sayı: ÖZEL SAYI

Serbest Sırada Birliktelik İstatistiklerinin Kullanımıyla Türkçe'nin Biçimbirimsel Belirsizliği'nin Giderilmesi

Morphological Disambiguation of Turkish with Free-order Co-occurrence Statistics

Pages
46–52
DOI
—

Abstract

Bu makalede, Türkçe gibi biçimsel olarak karmaşık yapıda olan dillerde sıklıkla karşılaşılan biçimbirimsel belirsizlik problemi için bir çözüm önerilmiştir. Genellikle, bu tipte bir problemin çözümü için bir cümledeki muhtemel kelime sıralarından uygun olanın seçilmesi için bilgiyi maksimuma çıkaran istatistiksel yöntemler uygulanmaktadır. Olasılıkların hesaplanması ve uygun sıranın seçilmesi için tercih edilecek metot uygulanacak dilin doğasına bağlıdır. Cümlelerde geçen kelimelerin madde başlarının oluşturduğu bir anlamsal çizgeden elde edilen birliktelik istatistikleri kullanılarak alternatifler arasından uygun olan kelime sıra dizilimi seçilmektedir. Bu çizge ağının belirsizlik içermeyen serbest sıralı karakteri istatistiklerin bağımsız olarak hesaplanmasında oldukça faydalıdır. Olasılıksal değerler Naive Bayes (NB) yöntemi kullanılarak elde edilmekte ve her kelime sıraları arasından uygun olanının, Viterbi algoritmasından esinlenilerek, maksimumu seçilmektedir.

Özet

In this article, a solution to the morphological ambiguity problem which occurs frequently in morphologically complex languages like Turkish is proposed. Generally, statistical methods are applicable for these tasks which maximize the information, obtained for a probable word order sequence in a sentence. The decision in selection of the method for calculation of the probabilities and the sequence selection method depends on the nature of the language. By using the co-occurrence statistics obtained from a semantic graph network which represents the lemmas of the sentences, the best word order sequence is selected from the alternatives. The non-ambiguous and free-word-order character of this network is helpful in determining the statistics independently. The probability values are obtained by using the Naive Bayes (NB) method and the selection of each word sequence is achieved by maximization, in the inspiration of the Viterbi algorithm.

Keywords: Co-occurrence, Morphological ambiguity, Naive Bayes, the Viterbi algorithm