HIRI-ViT: Scaling Vision Transformer With High Resolution Inputs.

Ting Yao,Yehao Li,Yingwei Pan,Tao Mei

doi:10.1109/tpami.2024.3379457

Abstract

The hybrid deep models of Vision Transformer (ViT) and Convolution Neural Network (CNN) have emerged as a powerful class of backbones for vision tasks. Scaling up the input resolution of such hybrid backbones naturally strengthes model capacity, but inevitably suffers from heavy computational cost that scales quadratically. Instead, we present a new hybrid backbone with HIgh-Resolution Inputs (namely HIRI-ViT), that upgrades prevalent four-stage ViT to five-stage ViT tailored for high-resolution inputs. HIRI-ViT is built upon the seminal idea of decomposing the typical CNN operations into two parallel CNN branches in a cost-efficient manner. One high-resolution branch directly takes primary high-resolution features as inputs, but uses less convolution operations. The other low-resolution branch first performs down-sampling and then utilizes more convolution operations over such low-resolution features. Experiments on both recognition task (ImageNet-1K dataset) and dense prediction tasks (COCO and ADE20 K datasets) demonstrate the superiority of HIRI-ViT. More remarkably, under comparable computational cost ( ∼5.0 GFLOPs), HIRI-ViT achieves to-date the best published Top-1 accuracy of 84.3% on ImageNet with 448×448 inputs, which absolutely improves 83.4% of iFormer-S by 0.9% with 224×224 inputs.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

HIRI-ViT: Scaling Vision Transformer With High Resolution Inputs.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on pattern analysis and machine intelligence

Lead the way for us

Journal: IEEE transactions on pattern analysis and machine intelligence	Publication Date: Sep 1, 2024
Citations: 5

Similar Papers

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
Wenhai Wang ... Enze Xie
-
Wenhai Wang, et. al.Wenhai Wang ... Enze Xie
01 Oct 2021
01 Oct 2021

Design and Implementation of a Deep Learning Target Detection System
...
-
, et. al. ...
30 Sep 2020
30 Sep 2020

Scalable Fire and Smoke Segmentation from Aerial Images Using Convolutional Neural Networks and Quad-Tree Search
Gonçalo Perrolas ... Alexandre Bernardino
Sensors | VOL. 22
Gonçalo Perrolas, et. al.Gonçalo Perrolas ... Alexandre Bernardino
22 Feb 2022
Sensors | VOL. 22

A Deep Convolutional Neural Network for Food Detection and Recognition
...
-
, et. al. ...
01 Dec 2018
01 Dec 2018

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

HIRI-ViT: Scaling Vision Transformer With High Resolution Inputs.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on pattern analysis and machine intelligence