VL-NMS: Breaking Proposal Bottlenecks in Two-stage Visual-language Matching

Chenchi Zhang,Wenbo Ma,Hanwang Zhang,Long Chen,Yueting Zhuang,Jun Xiao,Jian Shao

doi:10.1145/3579095

Abstract

The prevailing framework for matching multimodal inputs is based on a two-stage process: (1) detecting proposals with an object detector and (2) matching text queries with proposals. Existing two-stage solutions mostly focus on the matching step. In this article, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., query-agnostic), hoping that the proposals contain all instances mentioned in the text query (i.e., query-aware). Due to this mismatch, chances are that proposals relevant to the text query are suppressed during the filtering process, which in turn bounds the matching performance. To this end, we propose VL-NMS, which is the first method to yield query-aware proposals at the first stage. VL-NMS regards all mentioned instances as critical objects and introduces a lightweight module to predict a score for aligning each proposal with a critical object. These scores can guide the NMS operation to filter out proposals irrelevant to the text query, increasing the recall of critical objects, and resulting in a significantly improved matching performance. Since VL-NMS is agnostic to the matching step, it can be easily integrated into any state-of-the-art two-stage matching method. We validate the effectiveness of VL-NMS on three multimodal matching tasks, namely referring expression grounding, phrase grounding, and image-text matching. Extensive ablation studies on several baselines and benchmarks consistently demonstrate the superiority of VL-NMS.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

VL-NMS: Breaking Proposal Bottlenecks in Two-stage Visual-language Matching

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Multimedia Computing, Communications, and Applications

Lead the way for us

Similar Papers

Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding
Long Chen ... Wenbo Ma
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 35
Long Chen, et. al.Long Chen ... Wenbo Ma
18 May 2021
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 35

Two-stage matching decision-making method in medical service supply chain
Yan Qiu ... Yunxia Cao
International Journal of Logistics Research and Applications | VOL. 25
Yan Qiu, et. al.Yan Qiu ... Yunxia Cao
22 Jan 2021
International Journal of Logistics Research and Applications | VOL. 25

Hybrid Joint Embedding with Intra-Modality Loss for Image-Text Matching
Doaa B Ebaid ... Magda M Madbouly
-
Doaa B Ebaid, et. al.Doaa B Ebaid ... Magda M Madbouly
26 Nov 2022
26 Nov 2022

Cross-modal multi-relationship aware reasoning for image-text matching
Jin Zhang ... Luping Liu
Multimedia Tools and Applications | VOL. 81
Jin Zhang, et. al.Jin Zhang ... Luping Liu
27 Jan 2021
Multimedia Tools and Applications | VOL. 81

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

VL-NMS: Breaking Proposal Bottlenecks in Two-stage Visual-language Matching

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Multimedia Computing, Communications, and Applications