Vision Transformer with hierarchical structure and windows shifting for person re-identification.

Yinghua Zhang,Wei Hou

doi:10.1371/journal.pone.0287979

Yinghua Zhang, Wei Hou

Open Access

https://doi.org/10.1371/journal.pone.0287979

Copy DOI

Journal: PloS one	Publication Date: Jun 30, 2023
Citations: 1	License type: CC BY 4.0

Affiliation: Jiaozuo University, Henan University

Abstract

Extracting rich feature representations is a key challenge in person re-identification (Re-ID) tasks. However, traditional Convolutional Neural Networks (CNN) based methods could ignore a part of information when processing local regions of person images, which leads to incomplete feature extraction. To this end, this paper proposes a person Re-ID method based on vision Transformer with hierarchical structure and window shifting. When extracting person image features, the hierarchical Transformer model is constructed by introducing the hierarchical construction method commonly used in CNN. Then, considering the importance of local information of person images for complete feature extraction, the self-attention calculation is performed by shifting within the window region. Finally, experiments on three standard datasets demonstrate the effectiveness and superiority of the proposed method.

Full Text