Journals / Turkish Journal of Electrical Engineering and Computer Sciences / 2019 / Cilt: 27 - Sayı: 5

Sentence similarity using weighted path and similarity matrices

Pages
3779–3790
DOI
—

Abstract

Sentence similarity is the task of assessing how similar the two snippets of text are. Similarity techniques areused extensively in clustering, summarization, classification, plagiarism detection etc. Due to a small set of vocabularies,sentence similarity is considered to be a difficult problem in natural language processing. There are two issues in solvingthis problem: (1) Which similarity techniques to be used for word pair similarity and (2) How to generalize that tosentence pairs. We have used the weighted path, a WordNet-based similarity assessment, and the paraphrase databaseto obtain word pair similarity values. Thereafter, we extracted maximum values from the pairwise similarity matrixand computed a similarity value for a sentence pair. We have also incorporated a vector space model technique toform a robust similarity measure. Our method outperformed state-of-the-art methods on the STSS65 test dataset inPearson’s correlation of 87% compared to human similarity scores. Moreover, our approach performed on par with othermethods on the STSS131 test data using the same test. Our approach outperforms all the other WordNet-based methodscompared on both datasets.