MIA-Former: Efficient and Robust Vision Transformers via Multi-Grained Input-Adaptation

Zhongzhi Yu,Yingyan Lin,Sicheng Li,Yonggan Fu,Chaojian Li

doi:10.1609/aaai.v36i8.20879

Abstract

Vision transformers have recently demonstrated great success in various computer vision tasks, motivating a tremendously increased interest in their deployment into many real-world IoT applications. However, powerful ViTs are often too computationally expensive to be fitted onto real-world resource-constrained platforms, due to (1) their quadratically increased complexity with the number of input tokens and (2) their overparameterized self-attention heads and model depth. In parallel, different images are of varied complexity and their different regions can contain various levels of visual information, e.g., a sky background is not as informative as a foreground object in object classification tasks, indicating that treating those regions equally in terms of model complexity is unnecessary while such opportunities for trimming down ViTs' complexity have not been fully exploited. To this end, we propose a Multi-grained Input-Adaptive Vision Transformer framework dubbed MIA-Former that can input-adaptively adjust the structure of ViTs at three coarse-to-fine-grained granularities (i.e., model depth and the number of model heads/tokens). In particular, our MIA-Former adopts a low-cost network trained with a hybrid supervised and reinforcement learning method to skip the unnecessary layers, heads, and tokens in an input adaptive manner, reducing the overall computational cost. Furthermore, an interesting side effect of our MIA-Former is that its resulting ViTs are naturally equipped with improved robustness against adversarial attacks over their static counterparts, because MIA-Former's multi-grained dynamic control improves the model diversity similar to the effect of ensemble and thus increases the difficulty of adversarial attacks against all its sub-models. Extensive experiments and ablation studies validate that the proposed MIA-Former framework can (1) effectively allocate adaptive computation budgets to the difficulty of input images, achieving state-of-the-art (SOTA) accuracy-efficiency trade-offs, e.g., up to 16.5\% computation savings with the same or even a higher accuracy compared with the SOTA dynamic transformer models, and (2) boost ViTs' robustness accuracy under various adversarial attacks over their vanilla counterparts by 2.4\% and 3.0\%, respectively. Our code is available at https://github.com/RICE-EIC/MIA-Former.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MIA-Former: Efficient and Robust Vision Transformers via Multi-Grained Input-Adaptation

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Jun 28, 2022
Citations: 8

Similar Papers

Merry Go Round: Rotate a Frame and Fool a DNN
Daksh Thapar ... Chetan Arora
-
Daksh Thapar, et. al.Daksh Thapar ... Chetan Arora
01 Jun 2022
01 Jun 2022

DiFNet: Densely High-Frequency Convolutional Neural Networks
Wenzheng Hu ... Mingyang Li
IEEE Signal Processing Letters | VOL. 28
Wenzheng Hu, et. al.Wenzheng Hu ... Mingyang Li
01 Jan 2020
IEEE Signal Processing Letters | VOL. 28

Robustness-Aware Filter Pruning for Robust Neural Networks Against Adversarial Attacks
Hyuntak Lim ... Si-Dong Roh
-
Hyuntak Lim, et. al.Hyuntak Lim ... Si-Dong Roh
25 Oct 2021
25 Oct 2021

Adversarial Attack on Semantic Segmentation Preprocessed with Super Resolution
Gyeongsup Lim ... Junbeom Hur
-
Gyeongsup Lim, et. al.Gyeongsup Lim ... Junbeom Hur
21 Aug 2022
21 Aug 2022

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MIA-Former: Efficient and Robust Vision Transformers via Multi-Grained Input-Adaptation

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence