Journals / Avrupa Bilim ve Teknoloji Dergisi / 2020 / Cilt: 0 - Sayı: 20
Comparison of fastText and Bag of Words Word Representation Methods by Using Turkish Reviews Conducted for Touristic Places
- Pages
- 311–320
- DOI
- —
Abstract
Nowadays, with the increasing number and use of social media platforms, people now share their experiences about a product they havebought or a place they have been to on social media platforms more frequently. Considering the volume of data on social media platforms, it is considered that there is some meaningful information for institutions or companies in the reviews and experiences sharedon social media platforms. As such, it is important to improve the methods of extracting meaningful information from the reviews andexperiences shared on social media and to know which method is better. In this study, the classification successes of the bag of wordsand the fastText word representation methods, which are among the word representation methods in sentiment analysis methodsmentioned above, were compared by using Turkish reviews performed for touristic places. Besides, while performing the comparisonprocess, it was measured whether the process of separating the words into their roots and negation of the words, which is the preliminarystage of the sentiment analysis process, contributed to the classification success. In the study, both two-class (positive, negative)sentiment analysis and three-class (positive, negative, neutral) sentiment analysis were performed. Six data sets were created to carryout the mentioned comparison operations. The data sets were first classified using the Naive Bayes (NB), Multinomial Naive Bayes(MNB), k-Nearest Neighbor (k-NN) and Support Vector Machines (SVM) algorithms, which are frequently used in text mining, andbased on bag of words word representation method, they were classified with WEKA program. After the test results of all data sets wereobtained according to the bag of words word representation method, the tests of the fastText word representation method were carriedout using the fastText library of the Python programming language. Classification procedures were carried out with 10-fold cross validation methods, and f-score values of the classification processes were obtained. Finally, it was determined that bag of words wordrepresentation method performed a more successful classification than the fastText word representation method in two-class emotionanalysis, while the fastText word representation method performed a more successful classification process than bag of words wordrepresentation method in three-class emotional analysis. It was observed that the process of separating the words into their roots andnegating the words, which are the preliminary processes of sentiment analysis, did not contribute positively or negatively to theclassification processes performed with the fastText word representation method. However, it was determined that it had a minorcontribution to sentiment analysis processes performed by using bag of words word representation method. In the two-class sentimentanalysis, the most successful classification result was achieved by using the machine learning model created with the SVM algorithmwith the value of 0.91 f-score employing bag of words word representation method. In the three-class sentiment analysis, the mostsuccessful classification result was achieved with the machine learning model created using the fastText word representation methodwith the value of 0.78 f-score.
Özet
Günümüzde sosyal medya platformlarının sayısının ve kullanımının artmasıyla birlikte artık insanlar satın aldıkları bir ürünle veyagittikleri bir yer ile ilgili deneyimlerini sosyal medya platformlarında daha sıklıkla paylaşmaktadırlar. Sosyal medya platformlarındakiverilerin hacmi düşünüldüğünde, sosyal medya platformlarında paylaşılan incelemeler ve deneyimler içerisinde kurumlar veya şirketleriçin anlamlı birtakım bilgilerin olduğu düşünülmektedir. Hal böyle olunca sosyal medyada paylaşılan incelemeler ve deneyimler içerisinden anlamlı bilgi çıkarma yöntemlerini daha iyi hale getirmek ve hangi yöntemin daha iyi olduğunu bilmek önem arz etmektedir.Bu çalışmada turistik mekanlar için yapılan Türkçe incelemeler kullanılarak, yukarıda bahsedilen yöntemlerden biri olan duygu analiziyöntemindeki kelime temsil yöntemlerinden kelime çantası ve fastText kelime temsil yöntemlerinin sınıflandırma başarılarıkarşılaştırılmıştır. Ayrıca karşılaştırma işlemi gerçekleştirilirken duygu analizi işleminin ön hazırlık aşaması olan kelimeleri köklerineayırma ve kelimeleri olumsuzlaştırma işlemlerinin sınıflandırma başarısına katkılarının olup olmadığı ölçülmüştür. Çalışmada hem ikisınıflı (pozitif, negatif) duygu analizi hem de üç sınıflı (pozitif, negatif, nötr) duygu analizi gerçekleştirilmiştir. Bahsedilen karşılaştırmaişlemlerini gerçekleştirebilmek için altı adet veri seti oluşturulmuştur. Veri setleri önce metin madenciliğinde sıklıkla kullanılan NaiveBayes (NB), Multinom Naive Bayes (MNB), k-Nearest Neighbor (k-NN) ve Support Vector Machines (SVM) algoritmaları kullanılarakve kelime çantası kelime temsil yöntemi esas alınarak WEKA programıyla sınıflandırılmıştır. Tüm veri setlerinin kelime çantası kelimetemsil yöntemine göre test sonuçları elde edildikten sonra fastText kelime temsil yöntemine dair testler python programlama dilininfastText kütüphanesi kullanılarak gerçekleştirilmiştir. Sınıflandırma işlemleri 10 tekrarlı çapraz doğrulama yöntemiyle yapılaraksınıflandırma işlemlerinin f-skor değerleri elde edilmiştir. Nihayetinde iki sınıflı duygu analizinde kelime çantası kelime temsilyönteminin fastText kelime temsil yönteminden daha başarılı sınıflandırma gerçekleştirdiği, üç sınıflı duygu analizinde ise tam tersi bir şekilde fastText kelime temsil yönteminin kelime çantası kelime temsil yönteminden daha başarılı sınıflandırma işlemi gerçekleştirdiğitespit edilmiştir. Duygu analizi ön hazırlık işlemlerinden kelimeleri köklerine ayırma ve olumsuzlaştırma işlemlerinin fastText kelimetemsil yöntemiyle gerçekleştirilen sınıflandırma işlemlerinde olumlu ya da olumsuz bir katkı sağlamadığı görülmüştür. Ancak kelimeçantası kelime temsil yöntemi kullanılarak gerçekleştirilen duygu analizi işlemlerinde az da olsa bir katkısının olduğu tespit edilmiştir.İki sınıflı duygu analizinde en başarılı sınıflandırma sonucuna kelime çantası kelime temsil yöntemi kullanılarak 0.91 f-skoru değeriyleSVM algoritmasıyla oluşturulan makine öğrenmesi modeliyle ulaşılmıştır. Üç sınıflı duygu analizinde ise en başarılı sınıflandırmasonucuna 0.78 f-skoru değeriyle fastText kelime temsil yöntemi kullanılarak oluşturulan makine öğrenmesi modeliyle ulaşılmıştır.