Towards making the most of BERT in neural machine translation

Jian Yang ,Mingxuan Wang ,Hao Zhang ,Weinan Zhang ,Yu Yan ,Chengqi Zhao ,Lei Li

doi:10.5281/zenodo.6766335

Abstract

GPT-2 and BERT demonstrate the effectiveness of using pre-trained language models (LMs) on various natural language processing tasks. However, LM fine-tuning often suffers from catastrophic forgetting when applied to resource-rich tasks. In this work, we introduce a concerted training framework (CTNMT) that is the key to integrate the pre-trained LMs to neural machine translation (NMT). Our proposed CTNMT consists of three techniques: a) asymptotic distillation to ensure that the NMT model can retain the previous pre-trained knowledge; b) a dynamic switching gate to avoid catastrophic forgetting of pre-trained knowledge; and c) a strategy to adjust the learning paces according to a scheduled policy. Our experiments in machine translation show CTNMT gains of up to 3 BLEU score on the WMT14 English-German language pair which even surpasses the previous state-of-the-art pre-training aided NMT by 1.4 BLEU score. While for the large WMT14 English-French task with 40 millions of sentence-pairs, our base model still significantly improves upon the state-of-the-art Transformer big model by more than 1 BLEU score. The code and model can be downloaded from https://github.com/bytedance/neurst/ tree/master/examples/ctnmt.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Towards making the most of BERT in neural machine translation

Abstract

Talk to us

Similar Papers

More From: arXiv (Cornell University)

Lead the way for us

Similar Papers

Towards Making the Most of BERT in Neural Machine Translation
Jiacheng Yang ... Chengqi Zhao
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34
Jiacheng Yang, et. al.Jiacheng Yang ... Chengqi Zhao
03 Apr 2020
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34

Language Model Pre-training Method in Machine Translation Based on Named Entity Recognition
Zhen Li ... Chaojie Xie
International Journal on Artificial Intelligence Tools | VOL. 29
Zhen Li, et. al.Zhen Li ... Chaojie Xie
30 Nov 2020
International Journal on Artificial Intelligence Tools | VOL. 29

Adversarial Subword Regularization for Robust Neural Machine Translation
Jungsoo Park ... Jaewoo Kang
-
Jungsoo Park, et. al.Jungsoo Park ... Jaewoo Kang
01 Jan 2020
01 Jan 2020

On the Copying Behaviors of Pre-Training for Neural Machine Translation
...
-
, et. al. ...
01 Aug 2021
01 Aug 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Towards making the most of BERT in neural machine translation

Abstract

Talk to us

Similar Papers

More From: arXiv (Cornell University)