Journals / Turkish Journal of Electrical Engineering and Computer Sciences / 2020 / Cilt: 28 - Sayı: 1
An index-based joint multilingual/cross-lingual text categorization using topic expansion via BabelNet
- Pages
- 224–237
- DOI
- —
Abstract
: The majority of the state-of-the-art text categorization algorithms are supervised and therefore require priortraining. Besides the rigor involved in developing training datasets and the requirement for repetition of training fordifferent texts, working with multilingual texts poses additional unique challenges. One of these challenges is that thedeveloper is required to have many different languages involved. Term expansion such as query expansion has beenapplied in numerous applications; however, a major drawback of most of these applications is that the actual meaning ofterms is not usually taken into consideration. Considering the semantics of terms is necessary because of the polysemousnature of most natural language words. In this paper, as a specific contribution to the document index approach for textcategorization, we present a joint multilingual/cross-lingual text categorization algorithm (JointMC) based on semanticterm expansion of class topic terms through an optimized knowledge-based word sense disambiguation. The lexicalknowledge in BabelNet is used for the word sense disambiguation and expansion of the topics’ terms. The categorizationalgorithm computes the distributed semantic similarity between the expanded class topics and the text documents in thetest corpus. We evaluate our categorization algorithm using a multilabel text categorization problem. The multilabelcategorization task uses the JRC-Acquis dataset. The JRC-Acquis dataset is based on subject domain classification ofthe European Commission’s EuroVoc microthesaurus. We compare the performance of the classifier with a model ofit using the original class topics. Furthermore, we compare the performance of our classifier with two state-of-the-artsupervised algorithms (each for multilingual and cross-lingual tasks) using the same dataset. Empirical results obtainedon five experimental languages show that categorization with expanded topics shows a very wide performance marginwhen compared to usage of the original topics. Our algorithm outperforms the existing supervised technique, whichused the same dataset. Cross-language categorization surprisingly shows similar performance and is marginally betterfor some of the languages.