Journals / Eğitim Teknolojisi Kuram ve Uygulama / 2019 / Cilt: 9 - Sayı: 1
PREDICTING STUDENTS’ STEM CAREER INTERESTS BY USING DATA MINING APPROACH
- Pages
- 73–88
- DOI
- —
Abstract
In this study, it is aimed at creating a model that will predict whether secondary school students will continue their education and professional careers in an area related to STEM or not. Interaction dataset made available to the participants in ASSISTments Data Mining Competition 2017 is analyzed. This anonymized dataset consists of approximately 1 million click-stream records collected from 1709 students who used the intelligent tutoring system between 2004-2007. The dataset also contained a training dataset that includes information about whether 514 students in the dataset continued their STEM careers or not. For prediction, the performance of the Random Forest (RF), kNN, SVM (Support Vector Machine) and GMB (Generalized Regression Models Boosted) algorithms are compared. There was a class imbalance problem in training dataset, therefore, we compared various data balancing algorithms’ effect on the prediction algorithms. A 10-fold cross-validation was used to evaluate the performance of prediction models. As a result, the best performance was obtained when SVM algorithm and oversampling method were used together. In this case, the prediction model predicted over the students who prefer STEM careers with an accuracy of 66%. Features that are important while predicting STEM career preferences of students were also analyzed.
Özet
Bu çalışmada ortaokul öğrencilerinin, ASSISTments isimli zeki öğretim sistemindeki etkileşim verileri kullanılarak, eğitim ve mesleki kariyerlerine STEM ile ilgili bir alanda devam edip etmeyeceklerini tahmin edecek bir model oluşturulması amaçlanmıştır. Analizler 2017 yılında aynı amaçla düzenlenen ASSISTments Veri Madenciliği Yarışması’nda (ASSISTments Data Mining Competition 2017) katılımcılara sunulan veri seti ile gerçekleştirilmiştir. Veri seti, 2004-2007 yılları arasında sistemi kullanan 1709 öğrenciye ilişkin yaklaşık 1 milyon satırlık tıklama verisini içermektedir. Veriler, öğrencileri tanımlayan bilgiler silinerek katılımcılara sunulmuştur. Veri setinde 514 öğrencinin STEM kariyerine devam edip etmedikleri bilgisini içeren bir eğitim veri seti yer almaktadır. Tahmin modeli oluşturmak amacıyla Random Forest (RF), kNN, SVM (Support Vector Machine) ve GMB (Generalized Regression Models Boosted) algoritmaları kullanılmıştır. Veri setinde STEM tercih eden ve etmeyen öğrenciler arasında dengesiz dağılım bulunmaktadır. Bu nedenle farklı veri dengeleme yöntemlerinin modellerin tahmin performansına etkisi de test edilmiştir. Sonuçların değerlendirilmesi için 10-katlı çapraz geçerlilik yöntemi kullanılmıştır. Yapılan analizler sonucunda en iyi sınıflama performansına SVM algoritması ile yukarı örnekleme yönteminin birlikte kullanıldığı durumda ulaşılmıştır. Bu durumda oluşturulan tahmin modeli, STEM kariyeri tercih eden öğrencilerin %66’sını doğru olarak tahmin etmiştir. Aynı zamanda öğrencilerin STEM kariyer tercihlerini belirlemede önemli olan değişkenler de analiz edilmiştir.