Dergiler / Atmospheric Pollution Research / 2020 / Cilt: 11 - Sayı: 1

Application of k-means and hierarchical clustering techniques for analysis of air pollution: A review (1980–2019)

Sayfa
40–56
DOI
—

Özet

Clustering is an explorative data analysis technique used for investigating the underlying structure in the data. Itdescribed as the grouping of objects, where the objects share similar characteristics. Over the past 50 years,clustering has been widely applied to atmospheric science data in particular, climate and meteorological data.Since the 1980's, air pollution studies began employing clustering techniques, and has since been successful, andthe aim of this paper is to provide a review of such studies. In particular, two well known and commonly usedclustering methods i.e. k-means and hierarchical agglomerative, that have been applied in air pollution studieshave been reviewed. Air pollution data from two sources i.e. ground-based monitoring stations and air masstrajectories depicting pollutant pathways, have been included. Research works that have focused on spatiotemporalcharacteristics of air pollutants, pollutant behavior in terms of source, transport pathways, apportionmentand links to meteorological conditions, comprise much of the research works reviewed. A total of 100research articles were included during the period of 1980–2019. The purpose of the clustering approach, thespecific technique used and the data to which it was applied constitute much of the discussion presented in thisreview. Overall, the k-means technique has been extensively used among the studies, while average and Wardlinkages were the most frequently applied hierarchical clustering techniques. Reviews of clustering techniquesapplied in air pollution studies are currently lacking and this paper aims to fill that gap. In addition, and to thebest of the authors' knowledge, this is the first review dedicated to clustering applications in air pollution studies,and the first that covers the longest time span (1980–2019).