High-Order Interaction Learning for Image Captioning

Yanhui Wang,Wenhui Li,Yongdong Zhang,Ning Xu,An-An Liu

doi:10.1109/tcsvt.2021.3121062

Abstract

Image captioning aims at understanding various semantic concepts (e.g., objects and relationships) from an image and integrating them in a sentence-level description. Hence, it is necessary to learn the interaction among these concepts. If we define the context of the interaction to be involved in the <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">subject-predicate-object triplet, most current methods only focus on the single triplet for the first-order interaction to generate sentences. Intuitively, we humans are able to perceive the high-order interaction among concepts from two or more triplets to describe an image. For example, when we see the triplets <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">man-cutting-sandwich and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">man-with-knife , it is natural to integrate and predict the sentence <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">man cutting sandwich with knife . This depends on the high-order interaction between <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">cutting and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">knife in different triplets. Therefore, exploiting high-order interaction is expected to benefit image captioning and focus on reasoning. In this paper, we introduce the novel high-order interaction learning method over detected objects and relationships for image captioning under the umbrella of the encoder-decoder framework. We first extract a set of object and relationship features in an image. During the encoding stage, the interactive refining network is proposed to learn high-order representations by modeling intra- and inter-object feature interaction in the self-attention fashion. During the decoding stage, the interactive fusion network is proposed to integrate object and relationship information by strengthening their high-order interaction based on language context for sentence generation. In this way, we learn the object-relationship dependencies in different stages, which can provide abundant cues for both visual understanding and caption generation. Extensive experiments show that the proposed method can achieve competitive performances against the state-of-the-art methods on MSCOCO dataset. Additional ablation studies further validate its effectiveness.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

High-Order Interaction Learning for Image Captioning

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Circuits and Systems for Video Technology

Lead the way for us

Journal: IEEE Transactions on Circuits and Systems for Video Technology	Publication Date: Jul 1, 2022
Citations: 65

Similar Papers

Improving Image Captioning with Better Use of Caption
Zhan Shi ... Xiaodan Zhu
-
Zhan Shi, et. al.Zhan Shi ... Xiaodan Zhu
01 Jan 2020
01 Jan 2020

Synthesis of Vision and Language: Multifaceted Image Captioning Application
Arpit Gupta ... Ishita Kohli
INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT | VOL. 07
Arpit Gupta, et. al.Arpit Gupta ... Ishita Kohli
23 Dec 2023
INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT | VOL. 07

Image Captioning Based on Semantic Scenes.
Fengzhi Zhao ... Yi Lv
Entropy (Basel, Switzerland) | VOL. 26
Fengzhi Zhao, et. al.Fengzhi Zhao ... Yi Lv
18 Oct 2024
Entropy (Basel, Switzerland) | VOL. 26

Dynamic Convolution-based Encoder-Decoder Framework for Image Captioning in Hindi
Santosh Kumar Mishra ... Sriparna Saha
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 22
Santosh Kumar Mishra, et. al.Santosh Kumar Mishra ... Sriparna Saha
24 Mar 2023
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 22

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

High-Order Interaction Learning for Image Captioning

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Circuits and Systems for Video Technology