NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application

Chuhan Wu,Tao Qi,Yongfeng Huang,Qi Liu,Fangzhao Wu,Yang Yu

doi:10.18653/v1/2021.findings-emnlp.280

Abstract

Pre-trained language models (PLMs) like BERT have made great progress in NLP. News articles usually contain rich textual information, and PLMs have the potentials to enhance news text modeling for various intelligent news applications like news recommendation and retrieval. However, most existing PLMs are in huge size with hundreds of millions of parameters. Many online news applications need to serve millions of users with low latency tolerance, which poses huge challenges to incorporating PLMs in these scenarios. Knowledge distillation techniques can compress a large PLM into a much smaller one and meanwhile keeps good performance. However, existing language models are pre-trained and distilled on general corpus like Wikipedia, which has some gaps with the news domain and may be suboptimal for news intelligence. In this paper, we propose NewsBERT, which can distill PLMs for efficient and effective news intelligence. In our approach, we design a teacher-student joint learning and distillation framework to collaboratively learn both teacher and student models, where the student model can learn from the learning experience of the teacher model. In addition, we propose a momentum distillation method by incorporating the gradients of teacher model into the update of student model to better transfer useful knowledge learned by the teacher model. Extensive experiments on two real-world datasets with three tasks show that NewsBERT can effectively improve the model performance in various intelligent news applications with much smaller models.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application

Abstract

Talk to us

Similar Papers

Lead the way for us

Publication Date: Jan 1, 2021
Citations: 9	License type: cc-by

Similar Papers

NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application
...
-
, et. al. ...
23 Oct 2021
23 Oct 2021

NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application
-
-
--
21 Oct 2021
21 Oct 2021

On the Power of Pre-Trained Text Representations
Yu Meng ... Jiawei Han
-
Yu Meng, et. al.Yu Meng ... Jiawei Han
14 Aug 2021
14 Aug 2021

Neural Transfer Learning For Vietnamese Sentiment Analysis Using Pre-trained Contextual Language Models
An Pha Le ... Tran Vu Pham
-
An Pha Le, et. al.An Pha Le ... Tran Vu Pham
16 Dec 2021
16 Dec 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application

Abstract

Talk to us

Similar Papers