Dergiler / Bitlis Eren Üniversitesi Fen Bilimleri Dergisi / 2020 / Cilt: 9 - Sayı: 3

Gauss Karma Modellerin Özellikleri ve Değişken Parçalanmalarına Dayalı Kümeleme

Properties of Gaussian Mixture Models and Variable Segmentations Based Clustering

Sayfa
1377–1388
DOI
—

Özet

Bu çalışmada çok değişkenli verideki homojenlik ve heterojenlik durumları incelenmiş ve heterojen değişkenlerbelirlenmiştir. Değişkenlerdeki parçalanmaların (heterojenlik) normal karma dağılımlardaki bileşenlere denkgeldiği gösterilmiş ve alt grup sayıları belirlenmiştir. k-ortalamalar (k-means) algoritması ile değişkenlerdekiparçalanmalara atanan gözlemler belirlenmiş ve veri gruplandırma yapılmıştır. Değişkenlerdeki her birparçalanmanın Gauss Karma Modeldeki (GMM) bir kümelenmeye karşılık geldiği varsayımı altında muhtemelküme sayıları ve küme sayıları için aralık elde edilmiş ve küme sayılarına bağlı olarak model sayıları belirlenmiştir.Parçalanma (bileşen) sayısına bağlı model sayıları Genetik Algoritmalarla (GA) belirlenmiş ve En Çok OlabilirlikKestirimi (MLE) algoritması ile parametreler tahmin edilmiştir. Modele dayalı kümeleme yöntemi ile GaussKarma Modeller arasından veri yapısına uyan en iyi modelin seçimi log-olabilirlik, AIC ve BIC gibi bilgi kriterleriile belirlenmiştir.

Abstract

In this study, homogeneity and heterogeneity in multivariate data were examined and heterogeneous variables were determined. Partitions in the variables (heterogeneity) were shown to coincide with the components of normal mixed distributions and the number of subgroups was determined. K-means algorithms were used to determine the observations of the partition of the variables and the data was grouped. Under the assumption that each partition in the variables corresponds to a cluster in the Gaussian Mixture Model (GMM), the range for the possible number of clusters and clusters was obtained and the model numbers were determined based on the clusters. The number of models based on the number of components (partition) was determined by Genetic Algorithms (GA) and the parameters were estimated with Maximum Likelihood Estimation (MLE) algorithm. With the model based clustering method, the selection of the best model matching the data structure from the Gaussian Mixture Models was determined by information criteria such as log- olabilirlik, AIC and BIC.