Dergiler / Türkiye Bilişim Vakfı Bilgisayar Bilimleri ve Mühendisliği Dergisi / 2010 / Cilt: 3 Sayı: 1

Türkçe Dokümanlar İçin N-gram Tabanlı Yeni Bir Sınıflandırma(Ng-ind): Yazar, Tür ve Cinsiyet

Sayfa
11–19
DOI
—

Özet

Bu çalışmada Türkçe bir dokümanın türü, yazarı ve doküman yazarının cinsiyeti Türkçe’nin n-gram modeli kullanılarak belirlenmeye çalışılmıştır. N-gram modelinde 2-, 3-, 4-gram’lar kullanılmış ve üç farklı veri seti üzerinde toplam altı adet özellik vektörü oluşturulmuştur. Naive Bayes (NB), Destek Vektör Makinesi (DVM), Rastgele Orman (RO), K-En Yakın Komşuluk (K-EYK) gibi sınıflandırıcıların yanında geliştirdiğimiz Ng-ind yöntemi kullanılarak testler yapılmış ve başarı performansları birbirleri ile karşılaştırılmıştır. Ng-ind yöntemi cinsiyet ve tür belirlemede diğer yöntemlere göre daha iyi sonuç vermiştir. Bununla birlikte Ng-ind, tür belirlemede birleştirilmiş sınıflandırıcılardan da daha iyi performans göstermiştir.

Abstract

In this study, it is tried to find out a Turkish document’s genre, author and document author’s gender with using the Turkish n-gram model. In N-gram model, 2-, 3-, 4-grams were used, and total 6 feature vectors were produced on 3 different data set. Some tests were made with the Ng-ind method that we produced near the other classifiers such as Naive Bayes (NB), Support Vector Machine (SVM), Random Forest (RF), KNearest Neighbor (K-NN) and the success performances were compared with each other. In spite of the Ng-ind method gave better results than the other ones in gender and genre determination, it showed better performance than the compounded classifiers in genre determination