Siamese Alignment Network for Weakly Supervised Video Moment Retrieval

Yunxiao Wang,Yinwei Wei,Zhiyong Cheng,Liqiang Nie,Meng Liu,Yinglong Wang

doi:10.1109/tmm.2022.3168424

Abstract

Video moment retrieval, i.e., localizing the specific video moments within a video given a description query, has attracted substantial attention over the past several years. Although great progress has been achieved thus far, most of existing methods are supervised, which require moment-level temporal annotation information. In contrast, weakly-supervised methods which only need video-level annotations remain largely unexplored. In this paper, we propose a novel end-to-end Siamese alignment network for weakly-supervised video moment retrieval. To be specific, we design a multi-scale Siamese module, which could progressively reduce the semantic gap between the visual and textual modality with the Siamese structure. In addition, we present a context-aware multiple instance learning module by considering the influence of adjacent contexts, enhancing the moment-query and video-query alignment simultaneously. By promoting the matching of both moment-level and video-level, our model can effectively improve the retrieval performance, even if only having weak video level annotations. Extensive experiments on two benchmark datasets, i.e., ActivityNet-Captions and Charades-STA, verify the superiority of our model compared with several state-of-the-art baselines.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Siamese Alignment Network for Weakly Supervised Video Moment Retrieval

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia

Lead the way for us

Journal: IEEE Transactions on Multimedia	Publication Date: Jan 1, 2023
Citations: 14

Similar Papers

Word-Region Alignment-Guided Multimodal Neural Machine Translation
Yuting Zhao ... Chenhui Chu
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 30
Yuting Zhao, et. al.Yuting Zhao ... Chenhui Chu
01 Jan 2021
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 30

Multi-level textual-visual alignment and fusion network for multimodal aspect-based sentiment analysis
You Li ... Liang Chang
Artificial Intelligence Review | VOL. 57
You Li, et. al.You Li ... Liang Chang
01 Mar 2024
Artificial Intelligence Review | VOL. 57

Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing
Bingzhi Chen ... Zhongqi Wu
-
Bingzhi Chen, et. al.Bingzhi Chen ... Zhongqi Wu
01 Aug 2024
01 Aug 2024

Content-Based Video Retrieval With Temporal Localization Using a Deep Bimodal Fusion Approach
G. Megala ... P. Swarnalatha
-
G. Megala, et. al.G. Megala ... P. Swarnalatha
16 Jun 2023
16 Jun 2023

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Siamese Alignment Network for Weakly Supervised Video Moment Retrieval

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia