Towards Real-Time Panoptic Narrative Grounding by an End-to-End Grounding Network

Haowei Wang,Xiaoshuai Sun,Yongjian Wu,Jiayi Ji,Yiyi Zhou

doi:10.1609/aaai.v37i2.25350

Abstract

Panoptic Narrative Grounding (PNG) is an emerging cross-modal grounding task, which locates the target regions of an image corresponding to the text description. Existing approaches for PNG are mainly based on a two-stage paradigm, which is computationally expensive. In this paper, we propose a one-stage network for real-time PNG, termed End-to-End Panoptic Narrative Grounding network (EPNG), which directly generates masks for referents. Specifically, we propose two innovative designs, i.e., Locality-Perceptive Attention (LPA) and a bidirectional Semantic Alignment Loss (SAL), to properly handle the many-to-many relationship between textual expressions and visual objects. LPA embeds the local spatial priors into attention modeling, i.e., a pixel may belong to multiple masks at different scales, thereby improving segmentation. To help understand the complex semantic relationships, SAL proposes a bidirectional contrastive objective to regularize the semantic consistency inter modalities. Extensive experiments on the PNG benchmark dataset demonstrate the effectiveness and efficiency of our method. Compared to the single-stage baseline, our method achieves a significant improvement of up to 9.4% accuracy. More importantly, our EPNG is 10 times faster than the two-stage model. Meanwhile, the generalization ability of EPNG is also validated by zero-shot experiments on other grounding tasks. The source codes and trained models for all our experiments are publicly available at https://github.com/Mr-Neko/EPNG.git.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Towards Real-Time Panoptic Narrative Grounding by an End-to-End Grounding Network

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Jun 26, 2023
Citations: 4

Similar Papers

Sentence Generation for Entity Description with Content-Plan Attention
Bayu Trisedya ... Jianzhong Qi
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34
Bayu Trisedya, et. al.Bayu Trisedya ... Jianzhong Qi
03 Apr 2020
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 34

A small attentional YOLO model for landslide detection from satellite remote sensing images
Libo Cheng ... Mingguo Wang
Landslides | VOL. 18
Libo Cheng, et. al.Libo Cheng ... Mingguo Wang
22 May 2021
Landslides | VOL. 18

Towards Modeling Human Attention from Eye Movements for Neural Source Code Summarization
Aakash Bansal ... Bonita Sharif
Proceedings of the ACM on Human-Computer Interaction | VOL. 7
Aakash Bansal, et. al.Aakash Bansal ... Bonita Sharif
17 May 2023
Proceedings of the ACM on Human-Computer Interaction | VOL. 7

BADet: Boundary-Aware 3D Object Detection from Point Clouds
Rui Qian ... Xirong Li
Pattern Recognition | VOL. 125
Rui Qian, et. al.Rui Qian ... Xirong Li
10 Jan 2022
Pattern Recognition | VOL. 125

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Towards Real-Time Panoptic Narrative Grounding by an End-to-End Grounding Network

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence