Recognizing Transliterated English Words in Persian Texts

Ali Hoseinmardy,Saeedeh Momtazi

doi:10.29252/jist.8.30.84

Abstract

One of the most important problems of text processing systems is the word mismatch problem. This results in limited access to the required information in information retrieval. This problem occurs in analyzing textual data such as news, or low accuracy in text classification and clustering. In this case, if the text-processing engine does not use similar/related words in the same sense, it may not be able to guide you to the appropriate result. Various statistical techniques have been proposed to bridge the vocabulary gap problem; e.g., if two words are used in similar contexts frequently, they have similar/related meanings. Synonym and similar words, however, are only one of the categories of related words that are expected to be captured by statistical approaches. Another category of related words is the pair of an original word in one language and its transliteration from another language. This kind of related words is common in non-English languages. In non-English texts, instead of using the original word from the target language, the writer may borrow the English word and only transliterate it to the target language. Since this kind of writing style is used in limited texts, the frequency of transliterated words is not as high as original words. As a result, available corpus-based techniques are not able to capture their concept. In this article, we propose two different approaches to overcome this problem: (1) using neural network-based transliteration, (2) using available tools that are used for machine translation/transliteration, such as Google Translate and Behnevis. Our experiments on a dataset, which is provided for this purpose, shows that the combination of the two approaches can detect English words with 89.39% accuracy.

Highlights

Searching textual information on the Web has become one of the main usages of the Internet
To evaluate the performance of the proposed foreign word detection model, we provided a list of 979 words
Foreign words are English words that are transliterated from English to Persian; i.e., all words are written with Persian alphabets

Summary

Introduction

Searching textual information on the Web has become one of the main usages of the Internet. People use the Internet to find the information they need. For this reason, the intelligence of language processing tools can be much helpful in interacting with computers. One of the challenges encountered in text processing systems is recognizing related words in a language. The available methods, have not been able to produce good results for new words entered into a language. Transliterated words from dominated languages such as English to other languages are examples of new words in the target languages. Transliterated words are words that are entered in the target language with their vocals. These words are written in the target language as they are pronounced in the source language

Objectives

Results

Conclusion