Bottom-Up and Bidirectional Alignment for Referring Expression Comprehension

Liuwu Li,Yuqi Bu,Yi Cai

doi:10.1145/3474085.3475629

Abstract

In this paper, we propose a one-stage approach to improve referring expression comprehension (REC) which aims at grounding the referent according to a natural language expression. We observe that humans understand referring expressions through a fine-to-coarse bottom-up way, and bidirectionally obtain vision-language information between image and text. Inspired by this, we define the language granularity and the vision granularity. Otherwise, existing methods do not follow the mentioned way of human understanding in referring expression. Motivated by our observation and to address the limitations of existing methods, we propose a bottom-up and bidirectional alignment (BBA) framework. Our method constructs the cross-modal alignment starting from fine-grained representation to coarse-grained representation and bidirectionally obtains vision-language information between image and text. Based on the structure of BBA, we further propose a progressive visual attribute decomposing approach to decompose visual proposals into several independent spaces to enhance the bottom-up alignment framework. Experiments on five benchmark datasets of RefCOCO, RefCOCO+, ReferItGame, RefCOCOg and Flick30K show that our approach obtains +2.16%, +4.47%, +2.85%, +3.44%, and +2.91% improvements over the one-stage SOTA approaches, which validates the effectiveness of our approach.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Bottom-Up and Bidirectional Alignment for Referring Expression Comprehension

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point
Peizhi Zhao ... Yi Cai
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 38
Peizhi Zhao, et. al.Peizhi Zhao ... Yi Cai
24 Mar 2024
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 38

Multi-Task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
Gen Luo ... Chenglin Wu
-
Gen Luo, et. al.Gen Luo ... Chenglin Wu
01 Jun 2020
01 Jun 2020

Language-Attention Modular-Network for Relational Referring Expression Comprehension in Videos
Naina Dhingra ... Shipra Jain
-
Naina Dhingra, et. al.Naina Dhingra ... Shipra Jain
21 Aug 2022
21 Aug 2022

Continual Referring Expression Comprehension via Dual Modular Memorization.
Heng Tao Shen ... Cheng Chen
IEEE transactions on image processing : a publication of the IEEE Signal Processing Society | VOL. 31
Heng Tao Shen, et. al.Heng Tao Shen ... Cheng Chen
01 Jan 2021
IEEE transactions on image processing : a publication of the IEEE Signal Processing Society | VOL. 31

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Bottom-Up and Bidirectional Alignment for Referring Expression Comprehension

Abstract

Talk to us

Similar Papers