Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained Transformers

Minjia Zhang,Yuxiong He,Niranjan Uma Naresh

doi:10.1609/aaai.v36i10.21423

Abstract

Deep and large pre-trained language models (e.g., BERT, GPT-3) are state-of-the-art for various natural language processing tasks. However, the huge size of these models brings challenges to fine-tuning and online deployment due to latency and cost constraints. Existing knowledge distillation methods reduce the model size, but they may encounter difficulties transferring knowledge from the teacher model to the student model due to the limited data from the downstream tasks. In this work, we propose AD^2, a novel and effective data augmentation approach to improving the task-specific knowledge transfer when compressing large pre-trained transformer models. Different from prior methods, AD^2 performs distillation by using an enhanced training set that contains both original inputs and adversarially perturbed samples that mimic the output distribution from the teacher. Experimental results show that this method allows better transfer of knowledge from the teacher to the student during distillation, producing student models that retain 99.6\% accuracy of the teacher model while outperforming existing task-specific knowledge distillation baselines by 1.2 points on average over a variety of natural language understanding tasks. Moreover, compared with alternative data augmentation methods, such as text-editing-based approaches, AD^2 is up to 28 times faster while achieving comparable or higher accuracy. In addition, when AD^2 is combined with more advanced task-agnostic distillation, we can advance the state-of-the-art performance even more. On top of the encouraging performance, this paper also provides thorough ablation studies and analysis. The discovered interplay between KD and adversarial data augmentation for compressing pre-trained Transformers may further inspire more advanced KD algorithms for compressing even larger scale models.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained Transformers

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Jun 28, 2022
Citations: 4

Similar Papers

XtremeDistil: Multi-stage Distillation for Massive Multilingual Models
Subhabrata Mukherjee ... Ahmed Hassan Awadallah
-
Subhabrata Mukherjee, et. al.Subhabrata Mukherjee ... Ahmed Hassan Awadallah
01 Jan 2020
01 Jan 2020

Towards an Enhanced Understanding of Bias in Pre-trained Neural Language Models: A Survey with Special Emphasis on Affective Bias
Anoop K ... Lajish V L
-
Anoop K, et. al. Anoop K ... Lajish V L
01 Jan 2021
01 Jan 2021

Knowledge distillation and data augmentation for NLP light pre-trained models
Hanwen Luo ... Yuqing Zhang
Journal of Physics: Conference Series | VOL. 1651
Hanwen Luo, et. al.Hanwen Luo ... Yuqing Zhang
01 Nov 2020
Journal of Physics: Conference Series | VOL. 1651

AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
Daoyuan Chen ... Hongbo Deng
-
Daoyuan Chen, et. al.Daoyuan Chen ... Hongbo Deng
01 Jul 2020
01 Jul 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained Transformers

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence