Dergiler / Turkish Journal of Electrical Engineering and Computer Sciences / 2018 / Cilt: 26 - Sayı: 5
Enlarging multiword expression dataset by co-training
- Sayfa
- 2583–2594
- DOI
- —
Abstract
In multiword expressions (MWEs), multiple words unite to build a new unit in language. When MWEidentification is accepted as a binary classification task, one of the most important factors in performance is to train theclassifier with enough number of labelled samples. Since manual labelling is a time-consuming task, the performances ofMWE recognition studies are limited with the size of the training sets. In this study, we propose the comparison-basedand common-decision co-training approaches in order to enlarge the MWE dataset. In the experiments, the performancesof the proposed approaches were compared to those of the standard co-training [1] and manual labelling where statisticaland linguistic features are employed as two different views of the MWE dataset [2]. A number of tests with differentsettings were performed on a Turkish MWE dataset. Ten different classifiers were utilized in the experiments andthe best performing classifier pair was observed to be the SMO-SMO pair. The experimental results showed that thecommon-decision co-training approach is an alternative to hand-labeling of large MWE datasets and both newly proposedapproaches outperform the standard co-training [2] when the training set is to be enlarged in MWE classification.