Enhancing Multiscale Representations With Transformer for Remote Sensing Image Semantic Segmentation

Tao Xiao,Yuwen Huang,Yikun Liu,Gongping Yang,Mingsong Li

doi:10.1109/tgrs.2023.3256064

Abstract

Semantic segmentation is an extremely challenging task in high-resolution remote sensing (HRRS) images as objects have complex spatial layouts and enormous variations in appearance. Convolutional neural networks (CNNs) have excellent ability to extract local features and have been widely applied as the feature extractor for various vision tasks. However, due to the inherent inductive bias of convolution operation, CNNs inevitably have limitations in modeling long-range dependencies. Transformer can capture global representations well, but unfortunately ignores the details of local features and has high computational and spatial complexity in processing high-resolution feature maps. In this paper, we propose a novel hybrid architecture for HRRS image segmentation, termed EMRT, to exploit the advantages of convolution operations and Transformer to enhance multi-scale representation learning. We incorporate the deformable self-attention mechanism in the Transformer to automatically adjust the receptive field, and design an encoder-decoder architecture accordingly to achieve efficient context modeling. Specifically, the CNN is constructed to extract feature representations. In the encoder, local features and global representations at different resolutions are extracted by the CNN and Transformer, respectively, and fused in an interactive manner. Moreover, a separate spatial branch is designed to extract multi-scale contextual information as queries, and global dependencies between features at different scales are efficiently established by the decoder. Extensive experiments on three public remote sensing datasets demonstrate the superiority of EMRT and indicate that the overall performance of our method outperforms state-of-the-art methods. Code is available at https://github.com/peach-xiao/EMRT.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Enhancing Multiscale Representations With Transformer for Remote Sensing Image Semantic Segmentation

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Geoscience and Remote Sensing

Lead the way for us

Journal: IEEE Transactions on Geoscience and Remote Sensing	Publication Date: Jan 1, 2023
Citations: 25

Similar Papers

Self-constructing graph neural networks to model long-range pixel dependencies for semantic segmentation of remote sensing images
Qinghui Liu ... Arnt-Børre Salberg
International Journal of Remote Sensing | VOL. 42
Qinghui Liu, et. al.Qinghui Liu ... Arnt-Børre Salberg
16 Jun 2021
International Journal of Remote Sensing | VOL. 42

ResNet with Global and Local Image Features, Stacked Pooling Block, for Semantic Segmentation
Hui Song ... Xiaoqiang Guo
-
Hui Song, et. al.Hui Song ... Xiaoqiang Guo
01 Aug 2018
01 Aug 2018

CNN classification based on global and local features
Yufeng Zheng ... Matthias F Carlsohn
-
Yufeng Zheng, et. al.Yufeng Zheng ... Matthias F Carlsohn
14 May 2019
14 May 2019

Delving into deep representations for remote sensing image retrieval
Fan Hu ... Liangpei Zhang
-
Fan Hu, et. al.Fan Hu ... Liangpei Zhang
01 Nov 2016
01 Nov 2016

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Enhancing Multiscale Representations With Transformer for Remote Sensing Image Semantic Segmentation

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Geoscience and Remote Sensing