Dynamic Graph Attention for Referring Expression Comprehension

Sibei Yang,Guanbin Li,Yizhou Yu

doi:10.1109/iccv.2019.00474

Abstract

Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently requires visual reasoning on top of the relationships among the objects in the image. Meanwhile, the visual reasoning process is guided by the linguistic structure of the referring expression. However, existing approaches treat the objects in isolation or only explore the first-order relationships between objects without being aligned with the potential complexity of the expression. Thus it is hard for them to adapt to the grounding of complex referring expressions. In this paper, we explore the problem of referring expression comprehension from the perspective of language-driven visual reasoning, and propose a dynamic graph attention network to perform multi-step reasoning by modeling both the relationships among the objects in the image and the linguistic structure of the expression. In particular, we construct a graph for the image with the nodes and edges corresponding to the objects and their relationships respectively, propose a differential analyzer to predict a language-guided visual reasoning process, and perform stepwise reasoning on top of the graph to update the compound object representation at every node. Experimental results demonstrate that the proposed method can not only significantly surpass all existing state-of-the-art algorithms across three common benchmark datasets, but also generate interpretable visual evidences for stepwisely locating the objects referred to in complex language descriptions.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Dynamic Graph Attention for Referring Expression Comprehension

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Cops-Ref: A New Dataset and Task on Compositional Referring Expression Comprehension
Zhenfang Chen ... Lin Ma
-
Zhenfang Chen, et. al.Zhenfang Chen ... Lin Ma
01 Jun 2020
01 Jun 2020

One for all: One-stage referring expression comprehension with dynamic reasoning
Zhipeng Zhang ... Peng Wang
Neurocomputing | VOL. 518
Zhipeng Zhang, et. al.Zhipeng Zhang ... Peng Wang
13 Oct 2022
Neurocomputing | VOL. 518

Learning the Dynamics of Visual Relational Reasoning via Reinforced Path Routing
Chenchen Jing ... Qi Wu
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 36
Chenchen Jing, et. al.Chenchen Jing ... Qi Wu
28 Jun 2022
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 36

Referring Expression Comprehension Via Enhanced Cross-modal Graph Attention Networks
Jia Wang ... Wen-Huang Cheng
ACM Transactions on Multimedia Computing, Communications, and Applications | VOL. 19
Jia Wang, et. al.Jia Wang ... Wen-Huang Cheng
06 Feb 2023
ACM Transactions on Multimedia Computing, Communications, and Applications | VOL. 19

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Dynamic Graph Attention for Referring Expression Comprehension

Abstract

Talk to us

Similar Papers