Dergiler / İTÜ Dergisi Seri D: Mühendislik / 2007 / Cilt: 6 - Sayı: 5-6

Coğrafi veri seçim işlemi sonuçlarının değerlendirilmesinde hata matrisinin kullanımı

Usage of error matrix in the accuracy assessment of geographic data selection results

Sayfa
59–68
DOI
—

Özet

Bilgi sistemlerinin yoğun olarak kullanıldığı günümüzde, Coğrafi Bilgi Sistemlerine (CBS) ve CBS’nin temel bileşenlerinden coğrafi verilere olan ihtiyaç, artarak devam etmektedir. Coğrafi verilerin üretimi, CBS bileşenleri arasında en çok zaman ve maliyet gerektiren işlemi oluşturmaktadır. Dolayısıyla coğrafi veri üreticileri için, üretim süreçlerinin otomasyonu son derece önem taşımaktadır. İstenilen nitelik ve içerikteki coğrafi veriler, daha yüksek mekânsal çözünürlüğe sahip coğrafi verilerin bulunması durumunda bu verilerden türetme yolu ile üretilebilir. Türetilen coğrafi verilerin doğruluğu, CBS içindeki veriyi ve dolayısıyla analiz sonuçlarını doğrudan etkilediğinden, basılı haritalardaki durumlarından daha da önemli hale gelmektedir. Bilgisayar destekli harita üretim tekniklerin kullanımı ile birlikte, genelleştirme işlemleri sonuçlarının değerlendirilmesinde kullanılacak yeni yöntemlere olan gereksinim, kendini belirgin olarak göstermektedir. Hata matrisi, raster verilerle yapılan sınıflandırma işlemlerinin doğruluk analizlerinde yoğun olarak kullanılmaktadır. Bu çalışmada; hata matrisinin, türetme coğrafi veri üretim amaçlı gerçekleştirilen coğrafi veri seçim işlemi sonuçlarının değerlendirmesinde kullanımı irdelenmiştir. Türetme coğrafi verilerin üretiminde gerçekleştirilen coğrafi veri seçimi, aslında bir tür sınıflandırma işlemidir. Coğrafi detaylar, “seçilen” ve “seçilmeyen” şeklinde iki ayrı sınıfa ayrıştırılmaktadır. Hata matrisinin, coğrafi veri otomatik seçim sonuçlarının değerlendirilmesinde kullanımı bir uygulama ile gerçekleştirilmiştir. Hata matrisinin, yapılan seçim işlemi sonucu ve doğruluğu hakkında yansız ve kapsamlı bir değerlendirme sunduğu ve bu amaçla kullanılabileceği görülmüştür. Hata matrisinin, seçim işlemini doğrudan etkileyen en iyi seçim ölçütlerinin belirlenmesinde de kullanılabileceği değerlendirilmektedir.

Abstract

Nowadays, as being in the information age, the needs for Geographic Information Systems (GIS) and for its main component geographic data are incrementally increasing. Geographic data is the most expensive and time consuming component of GIS (Longley et al., 2001). Therefore construction of automation processes in geographic data production has high importance for data providers. In case of the availability of data having higher spatial resolution, derived geographic data with certain contents and specifications can be produced through generalization. Generalization is important to GIS, not only because of to drive towards automating the process to obtain derived databases, but also because generalization can seriously affect data within a GIS, even more so than within paper maps (Joao, 1998). A fully automated generalization cannot be anticipated at this moment (Spiess et al., 2005). Still lots of efforts are needed to automate the generalization processes (Stoter, 2005). And new automated methods need new quality control methods (Jaakola, 1994; Ruas, 2001). In this study, the usage of error (confusion) matrix is investigated in the accuracy assessment of selection process aiming to obtain derived geographic data.Geographic data selection process constitutes an important step in the production of derived datasets through generalization. The results of the selection process build the main skeleton of the derived data. Selection process and criteria used are highly depended on the intended usage purpose of the generalized data. In the selection process, unwanted features are eliminated and important features are retained. It is the one of the most immediate differences between maps of the same region at different scales (Joao, 1998). Contextual relations among geographic features should be considered in the selection process. Although selection process of geographic features constitutes an important step in the production of derived datasets, it is not only enough by itself. It should be considered together with other generalization operators (Shea and McMaster, 1989) in map or geographic data production line. Error matrix is generally used in the accuracy assessments of supervised and unsupervised classification of raster data (Congalton ve Green, 1999). An error matrix is a square array of numbers set out in rows and columns which expresses the number of sample units assigned to a particular class relative to actual class as verified by some reference data (Figure 3). The selection processes in the production of derived geographic data can be assessed as a kind of classification. Geographic data is classified in two different categories as “selected” or “eliminated”. The usage of error matrix in the accuracy assessment of selection results has been realized with an application. In the application, the automatic selection result of running waters and canals for the production of derived topographic maps has been used.Geographic data covered by Balıkesir J19 topographic map sheet was chosen as a test data. 1:25.000 scale content data was used for the automatic selection of running waters and canals for the production of 1:100.000 scale Balıkesir J19 topographic map sheet. Last edition of 1:100.000 scale Balıkesir J19 map sheet has been used as a reference data that is needed for verifying the selection result and constructing the error matrix. Due to the subjectivity of generalization, extra comments have been made in defining the accuracy of the selection results of individual features, in case of conflicts occurred between the selection results and the reference data. The lengths of geographic features can also be used in constructing error matrix. In this study, two types of error matrixes were constructed, using feature numbers and feature lengths, respectively. Then the accuracies of classification as a whole and as individual categories have been computed by using the error matrixes constructed. As a result, it has been found that error matrix provides objective and thorough information for the accuracy assessment of the selection process results for obtaining derived geographic data. It can be used for this purpose. Selection criteria would directly affect the selection results. Determining optimum criteria that will be used in the production lines constitute a critical issue. The Kappa coefficient of agreement (Cohen, 1960), which takes into account all the elements of the error matrix, yields a robust assessment for the accuracy of whole classification process and can be used in determining the optimum selection criteria.