Aligning Image Semantics and Label Concepts for Image Multi-Label Classification

Wei Zhou,Peng Dou,Tao Su,Haifeng Hu,Zhiwu Xia

doi:10.1145/3550278

Abstract

Image multi-label classification task is mainly to correctly predict multiple object categories in the images. To capture the correlation between labels, graph convolution network based methods have to manually count the label co-occurrence probability from training data to construct a pre-defined graph as the input of graph network, which is inflexible and may degrade model generalizability. Moreover, most of the current methods cannot effectively align the learned salient object features with the label concepts, so that the predicted results of model may not be consistent with the image content. Therefore, how to learn the salient semantic features of images and capture the correlation between labels, and then effectively align them is one of the key to improve the performance of image multi-label classification task. To this end, we propose a novel image multi-label classification framework which aims to align I mage S emantics with L abel C oncepts ( ISLC ). Specifically, we propose a residual encoder to learn salient object features in the images, and exploit the self-attention layer in aligned decoder to automatically capture the correlation between labels. Then, we leverage the cross-attention layers in aligned decoder to align image semantic features with label concepts, so as to make the labels predicted by model more consistent with image content. Finally, the output features of the last layer of residual encoder and aligned decoder are fused to obtain the final output feature for classification. The proposed ISLC model achieves good performance on various prevalent multi-label image datasets such as MS-COCO 2014, PASCAL VOC 2007, VG-500, and NUS-WIDE with 87.2%, 96.9%, 39.4%, and 64.2%, respectively.

Full Text

Published version (

Free)

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Aligning Image Semantics and Label Concepts for Image Multi-Label Classification

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Multimedia Computing, Communications, and Applications

Lead the way for us

Journal: ACM Transactions on Multimedia Computing, Communications, and Applications	Publication Date: Feb 6, 2023
Citations: 5

Similar Papers

Two‐level salient feature complementary network for person re‐identification
Haishun Du ... Zhaoyang Li
International Journal of Intelligent Systems | VOL. 37
Haishun Du, et. al.Haishun Du ... Zhaoyang Li
14 Jan 2022
International Journal of Intelligent Systems | VOL. 37

A CBIR Scheme Using GLCM Features in DCT Domain
Sumit Kumar ... Arun Kumar Pal
-
Sumit Kumar, et. al.Sumit Kumar ... Arun Kumar Pal
01 Dec 2017
01 Dec 2017

Visual Interaction Perceptual Network for Blind Image Quality Assessment
Xiaoqi Wang ... Weisi Lin
IEEE Transactions on Multimedia | VOL. 25
Xiaoqi Wang, et. al.Xiaoqi Wang ... Weisi Lin
01 Jan 2023
IEEE Transactions on Multimedia | VOL. 25

Scene semantics involuntarily guide attention during visual search.
Taylor R Hayes ... John M Henderson
Psychonomic Bulletin & Review | VOL. 26
Taylor R Hayes, et. al.Taylor R Hayes ... John M Henderson
24 Jul 2019
Psychonomic Bulletin & Review | VOL. 26

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Aligning Image Semantics and Label Concepts for Image Multi-Label Classification

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Multimedia Computing, Communications, and Applications