Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering

Jianwen Jiang,Yue Gao,Ziqiang Chen,Haojie Lin,Xibin Zhao

doi:10.1609/aaai.v34i07.6766

Abstract

Understanding questions and finding clues for answers are the key for video question answering. Compared with image question answering, video question answering (Video QA) requires to find the clues accurately on both spatial and temporal dimension simultaneously, and thus is more challenging. However, the relationship between spatio-temporal information and question still has not been well utilized in most existing methods for Video QA. To tackle this problem, we propose a Question-Guided Spatio-Temporal Contextual Attention Network (QueST) method. In QueST, we divide the semantic features generated from question into two separate parts: the spatial part and the temporal part, respectively guiding the process of constructing the contextual attention on spatial and temporal dimension. Under the guidance of the corresponding contextual attention, visual features can be better exploited on both spatial and temporal dimensions. To evaluate the effectiveness of the proposed method, experiments are conducted on TGIF-QA dataset, MSRVTT-QA dataset and MSVD-QA dataset. Experimental results and comparisons with the state-of-the-art methods have shown that our method can achieve superior performance.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Apr 3, 2020
Citations: 71

Similar Papers

Video Question Answering Using a Forget Memory Network
Yuanyuan Ge ... Youjiang Xu
-
Yuanyuan Ge, et. al.Yuanyuan Ge ... Youjiang Xu
01 Jan 2017
01 Jan 2017

Video question answering via grounded cross-attention network learning
Yunan Ye ... Jun Xiao
Information Processing & Management | VOL. 57
Yunan Ye, et. al.Yunan Ye ... Jun Xiao
16 Apr 2020
Information Processing & Management | VOL. 57

Spatio-Temporal Context Networks for Video Question Answering
Kun Gao ... Yahong Han
-
Kun Gao, et. al.Kun Gao ... Yahong Han
01 Jan 2018
01 Jan 2018

Hierarchical Conditional Relation Networks for Multimodal Video Question Answering
Thao Minh Le ... Svetha Venkatesh
International Journal of Computer Vision | VOL. 129
Thao Minh Le, et. al.Thao Minh Le ... Svetha Venkatesh
27 Aug 2021
International Journal of Computer Vision | VOL. 129

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence