Machine Translation of Mathematical Text

Aditya Ohri,Tanya Schmah

doi:10.1109/access.2021.3063715

Abstract

We have implemented a machine translation system, the PolyMath Translator, for LaTeX documents containing mathematical text. The current implementation translates English LaTeX to French LaTeX, attaining a BLEU score of 53.6 on a held-out test corpus of mathematical sentences. It produces LaTeX documents that can be compiled to PDF without further editing. The system first converts the body of an input LaTeX document into English sentences containing math tokens, using the pandoc universal document converter to parse LaTeX input. We have trained a Transformer-based translator model, using OpenNMT, on a combined corpus containing a small proportion of domain-specific sentences. Our full system uses this Transformer model and also Google Translate with a custom glossary, the latter being used as a backup to better handle linguistic features that do not appear in our training dataset. Google Translate is used when the Transformer model does not have confidence in its translation, as determined by a high perplexity score. Ablation testing demonstrates that the tokenization of symbolic expressions is essential to the high quality of translations produced by our system. We have published our test corpus of mathematical text. The PolyMath Translator is available as a web service at http://www.polymathtrans.ai.

Highlights

Machine translation for specialized domains such as legal or medical text has received considerable attention
We examine how mathematical text differs from text in other domains, and provide evidence that mathematical text is simpler in some respects
We evaluate our translation system on both whole LATEX documents and a small test corpus of mathematical sentence pairs, and perform ablation testing to evaluate the importance of different components of our system

Summary

Introduction

Machine translation for specialized domains such as legal or medical text has received considerable attention. Advances in these areas have been useful in practice and have given rise to new techniques in areas including domain adaptation [4], automatic term extraction [25] and domain-aware approaches to general-purpose machine translation [2]. The domain of mathematical text has, to our knowledge, not yet been the subject of research in machine translation, beyond some very early work [14] [17], we mention extensive ongoing research in the related areas of mathematical ontology and semantics [31], translation from informal mathematical writing into formal mathematics [29], and mathematical information retrieval [10]

Objectives

Methods

Results

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEEE Access	Publication Date: Jan 1, 2021
Citations: 25	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Machine Translation of Mathematical Text

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: IEEE Access

Lead the way for us

Similar Papers

Improving neural machine translation with POS-tag features for low-resource language pairs
Zar Zar Hlaing ... Ponrudee Netisopakul
Heliyon | VOL. 8
Zar Zar Hlaing, et. al.Zar Zar Hlaing ... Ponrudee Netisopakul
01 Aug 2022
Heliyon | VOL. 8

Automated and Human Interaction in Written Discourse: A Contrastive Parallel Corpus-based Investigation of Metadiscourse Features in Machine-Human Translations
Muhammad Afzaal ... Muhammad Imran
SAGE Open | VOL. 12
Muhammad Afzaal, et. al.Muhammad Afzaal ... Muhammad Imran
01 Oct 2022
SAGE Open | VOL. 12

Adapting transformer-based language models for heart disease detection and risk factors extraction
Essam H Houssein ... Abdelmgeid A Ali
Journal of Big Data | VOL. 11
Essam H Houssein, et. al.Essam H Houssein ... Abdelmgeid A Ali
04 Apr 2024
Journal of Big Data | VOL. 11

Google and Legal Translation: The Case Study of Contracts
Noor Riyadh Rahim
Arab World English Journal For Translation and Literary Studies | VOL. 8
Noor Riyadh RahimNoor Riyadh Rahim
26 May 2024
Arab World English Journal For Translation and Literary Studies | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Machine Translation of Mathematical Text

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: IEEE Access