MINIMAL: Mining Models for Universal Adversarial Triggers

Yaman Kumar Singla,Rajiv Ratn Shah,Balaji Krishnamurthy,Changyou Chen,Somesh Singh,Swapnil Parekh

doi:10.1609/aaai.v36i10.21384

Abstract

It is well known that natural language models are vulnerable to adversarial attacks, which are mostly input-specific in nature. Recently, it has been shown that there also exist input-agnostic attacks in NLP models, called universal adversarial triggers. However, existing methods to craft universal triggers are data intensive. They require large amounts of data samples to generate adversarial triggers, which are typically inaccessible by attackers. For instance, previous works take 3000 data samples per class for the SNLI dataset to generate adversarial triggers. In this paper, we present a novel data-free approach, MINIMAL, to mine input-agnostic adversarial triggers from models. Using the triggers produced with our data-free algorithm, we reduce the accuracy of Stanford Sentiment Treebank’s positive class from 93.6% to 9.6%. Similarly, for the Stanford Natural LanguageInference (SNLI), our single-word trigger reduces the accuracy of the entailment class from 90.95% to less than 0.6%. Despite being completely data-free, we get equivalent accuracy drops as data-dependent methods

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MINIMAL: Mining Models for Universal Adversarial Triggers

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Jun 28, 2022
Citations: 1

Similar Papers

Generating Label Cohesive and Well-Formed Adversarial Claims
Pepa Atanasova ... Dustin Wright
-
Pepa Atanasova, et. al.Pepa Atanasova ... Dustin Wright
01 Jan 2020
01 Jan 2020

Dynamic Cascades with Bidirectional Bootstrapping for Action Unit Detection in Spontaneous Facial Behavior.
Yunfeng Zhu ... Jeffrey F Cohn
IEEE transactions on affective computing | VOL. 2
Yunfeng Zhu, et. al.Yunfeng Zhu ... Jeffrey F Cohn
01 Apr 2011
IEEE transactions on affective computing | VOL. 2

DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
Tao Shen ... Shirui Pan
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 32
Tao Shen, et. al.Tao Shen ... Shirui Pan
27 Apr 2018
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 32

Developing of NLP models by model based software developement
Zsolt Krutilla
Gradus | VOL. 10
Zsolt KrutillaZsolt Krutilla
01 Jan 2023
Gradus | VOL. 10

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MINIMAL: Mining Models for Universal Adversarial Triggers

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence