Hybrid approach for semantic similarity calculation between Tamil words

Deepa Karuppaiah,Durai Raj Vincent P M

doi:10.1504/ijica.2021.10027869

Abstract

Semantic similarity, sometimes referred as semantic relatedness, is one of the important concepts that help in various applications that involve natural language processing. In literature, there are plenty of similarity measures to compute the relationship among words in monolingual and cross-lingual documents. They help us in understanding text, finding plagiarism, information retrieval etc. They can be categorised based on the resources used into corpus-based and knowledge-based measures. These measures are plenty for the English language. For the Tamil language, there are hardly any works in calculating the similarity between words. In this paper, we proposed a similarity finding technique that exploits the knowledge from the resources like Tamil Indo WordNet, Tamil Wikitionary and Oxford Tamil Dictionary. We have used the definitions and example sentences of each word that are available through each of these resources for similarity calculation. The proposed approach is evaluated using human evaluated Miller Charles and Rubenstein Goodenough datasets.

Full Text