Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering.

Zhou Yu,Dacheng Tao,Jianping Fan,Chenchao Xiang,Jun Yu

doi:10.1109/tnnls.2018.2817340

Abstract

Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering.

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Neural Networks and Learning Systems

Lead the way for us

Journal: IEEE Transactions on Neural Networks and Learning Systems	Publication Date: Apr 9, 2018
Citations: 486

Similar Papers

Visual Question Answering as Reading Comprehension
Hui Li ... Anton Van Den Hengel
-
Hui Li, et. al.Hui Li ... Anton Van Den Hengel
01 Jun 2019
01 Jun 2019

Multimodal feature fusion by relational reasoning and attention for visual question answering
Weifeng Zhang ... Zengchang Qin
Information Fusion | VOL. 55
Weifeng Zhang, et. al.Weifeng Zhang ... Zengchang Qin
19 Aug 2019
Information Fusion | VOL. 55

Multi-modal Factorized Bilinear Pooling with Co-attention Learning for Visual Question Answering
Zhou Yu ... Dacheng Tao
-
Zhou Yu, et. al.Zhou Yu ... Dacheng Tao
01 Oct 2017
01 Oct 2017

An Enhanced Term Weighted Question Embedding for Visual Question Answering
Sruthy Manmadhan ... Binsu C Kovoor
Journal of Information & Knowledge Management | VOL. 21
Sruthy Manmadhan, et. al.Sruthy Manmadhan ... Binsu C Kovoor
05 May 2022
Journal of Information & Knowledge Management | VOL. 21

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering.

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Neural Networks and Learning Systems