Dergiler / İTÜ Dergisi Seri C: Fen Bilimleri / 2006 / Cilt: 4 - Sayı: 1

Çok yönlü frekans tablolarının analizi üzerine bir çalışma

A study on analysis of multiway frequency tables

Sayfa
17–27
DOI
—

Özet

Çok yönlü frekans tablolarının analizi, iki veya daha fazla düzeyi olan üç veya daha fazla kesikli bağımsız değişkenler arasındaki ilişkileri değerlendirmek amacıyla kullanılan parametrik olmayan bir testtir. Bu değişkenler, sınıflama, nicel veya kategorik olabilir. Çok yönlü frekans analizi tablosunda, bağımlı değişken, bir ya da daha fazla kesikli değişkenler ve ilişkileri tarafından etkilenen hücre frekansıdır. Bu analiz, değişkenlerden biri bağımlı değişken olduğu zaman kullanılır. Bu durumda, temel etkiler ve etkileşimler test edilecektir. Bu analizde, ilk olarak etkiler tanımlanacak ve böylece tüm etkilerin her bir düzeyi için parametre tahminleri elde edilecektir. Eğer, kesikli ve sürekli değişkenlerin bir karışımı kullanılmak istenirse, genellikle Lojistik regresyon seçilir. Çok yönlü frekans analizi, Log-lineer analiz ve Lojit modellerin bir çeşididir. Bu üç analiz çeşidi, uyum iyiliği için Ki-karenin bir uzantısıdır. Elimizde var olan frekans verisi için hesaplama türleri, en fazla iki boyutlu olumsallık tabloları için uygundur. Böyle iki boyutlu tablolar, uyum iyiliği yaklaşımı ile analiz edilir. Çok yönlü frekans analizinde, Pearson Ki-kare yerine olabilirlik oranG2 istatistiği kullanılacaktır çünkü, Pearson Ki-kare istatistiği toplamsal değil iken, iç-içe modeller için bu test toplamsaldır. Ayrıca, bu çalışmada kısmi birliktelik testi, modelde her bir bireysel etkinin anlamlılığını test etmek için kullanılır. Bu çalışmanın amacı, kesikli değişkenler arasında bir ilişki olup olmadığını araştırmak, en iyi modeli belirlemek ve uygun modele ilişkin parametre tahminlerini elde etmektir. Bunun için, Ankara ilinde devlet tiyatrolarının izleyici profilini çıkartmak amacıyla yapılmış bir anket çalışmasından yararlanılmıştır.

Abstract

Multiway frequency analysis is a nonparametric test that can be used to evaluate relations among three or more discrete independent variables with two or more levels. These variables may be nominal or categorical as qualitative. In a multiway frequency analysis table, cell frequency is the dependent variable that is influenced by one or more discrete variables and their associations. Multiway frequency analysis can be used when one of the variables is a dependent variable. In that case, main effects and interactions can be tested. In this analysis, effects are identified first and then parameter estimates are found for each level of all the effects. If the researcher wishes to use a mix of continuous and discrete variables, Logistic regression is usually the method of choice. Multiway frequency analysis is a version of Loglinear analysis and Logit models. These three analysis version are extensions of the chi-square for goodness-of-fit. The computational techniques for handling frequency data are appropriate for contingency tables of at most two dimensions. Such two-dimensional tables are analysed with the goodness-of-fit approach. In multiway frequency analysis, we use the Likelihood ratio $G^2$ statistic instead of the Pearson Chi-square because it is additive for nested models, whereas the Pearson statistic, in general, is not. The goal of this study is to discover whether there is an association among discrete variables, to determine the best model and to obtain parameter estimates of the derived model. Public survey which has been done for the evaluation of the profile of governmental theatre attendees in Ankara were used for this purpose. Likelihood ratio $G^2$ statistic has the desirable property that it is additive, meaning the sum of the chisquare values for the individual effects in the model equals the chi-square for the total model. Therefore, if one considers the difference between two Likelihood chi-square statistics for related models, the result would be another Likelihood ratio $G^2$ statistic. This property enables one to make two important inferences: 1) nested models can be compared 2) individual effects may be assessed Prior to proceeding with model selection, a special case of the Chi-square test, Likelihood ratio $G^2$ statistic is performed. With $g_f$ as the observed frequency and $B_f$ as the expected frequency for each of k cells, the Pearson chi-square is computed as: $chi^2=sumlimits_{f=1}^{k}frac{(B_f-g_f)^2}{g_f}$Large differences yields large values of the statistics and more evidence that the model is inadequate. The Likelihood ratio $G_2$ statistic is computed by the following equation. $G^2 =2Sigma g_fLn(g_f/B_f)$ $G^2$ was developed by Fisher (1924), based on his earlier work on maximum likelihood theory. As with $X^2$, $G^2$ is distributed as $X^2$ for sufficiently large N. In this study, before testing and selecting a model that best fits the observed data, screening is normally carried out to see if there are any significant effects to investigate. The researcher may be specifically interested in finding the best model is developed where an additive regression type equation is written for expected frequency as a function of the effects in the design. The procedure is similar to multiple regressions except that cell frequencies are predicted based on the combined effects of indepent variables. The modelling process begins with all possible associations among the independent variables. That is, if there are three independent variable, it will start off with all of the one-, two-, and three way associations. This is called a full or saturated model, because it includes all possible effects. It has the same amount of cells in the frequency table as it does effects, so the expected cell frequencies will always exactly match the observed frequencies. The statistic normally used to tell you how well the model of expected frequencies fits the observed frequencies is the Likelihood ratio $G^2$statistic. In this study, partial associations test allow us to test the significance of each individual effect in the model. However, the likelihood ratio $G^2$ statistic can be used to compare a more complicated model with a simpler model which has one interaction or main effect dropped to assess the importance of that term.