Journals / Pegem Eğitim ve Öğretim Dergisi / 2019 / Cilt: 9 - Sayı: 4
Veri madenciliğinde kullanılan karar ağaçlarının karşılaştırılması
- Pages
- 1183–1208
- DOI
- —
Abstract
Bu çalışmanın amacı son yıllarda farklı alanlarda kullanılan veri madenciliği yöntemleri tarafından elde edilen karar ağaçlarının farklı ölçütlere göre karşılaştırılmasıdır. Çalışmada PISA 2015 öğrenci anketinde yer alan 12 bağımsız değişken yardımıyla öğrencileri fen okuryazarlığı bakımından başarılı ve başarısız olarak sınıflama amacıyla farklı yöntemler tarafından elde edilen karar ağaçlarının benzer ve farklı yönleri ortaya çıkarılmıştır. Türkiye örnekleminde 15 yaş grubundaki toplam 5895 öğrenciden elde edilen veriler Java tabanlı ve açık kaynak kodlu WEKA programında analiz edilmiştir. Analiz sonucunda doğru sınıflama oranları bakımından en başarılı yöntemlerin sırasıyla lojistik model, Hoefding tree, J.48, REPTree ve Random tree olduğu belirlenmiştir. Bunun yanında faklı öğrenme yöntemleri tarafından elde edilen karar ağaçlarında sınıflamada etkili olan değişkenlerin farklılık gösterdiği belirlenmiştir. Elde edilen sonuçlara göre farklı yöntemler tarafından elde edilen karar ağaçlarında öğrencileri sınıflamada etkili olan bağımsız değişkenlerin farklılık gösterdiği ve tek bir yöntem yerine birden fazla yönteme ilişkin bulguların rapor edilmesi önerilmiştir.
Özet
The purpose of this study is to compare decision trees obtained by data mining algorithms used in various areas in recent years according to different criteria. In the study, similar and different aspects of the decision trees obtained by different methods for classifying the students as successful and unsuccessful in terms of science literacy were revealed with the help of 12 independent variables included in the PISA 2015 student survey. Data collected across Turkey, from a total of 5895 students in the age group of 15, were analyzed in Java-based Weka software, which has an open source code. As a result of the analysis, it was found that the most successful algorithms in terms of correct classification rate were respectively Logistic Model, Hoeffding Tree, J.48, REPTree and Random Tree. In addition, regarding the decision trees obtained by different learning algorithms, variables that have been effective in the classification were found to be different. According to the results, it was concluded that independent variables found to be effective in the classification of the students for the decision trees obtained by different algorithms differed from each other and it was suggested to report the finding of more than one algorithm instead of those of only one algorithm.