Accuracy vs. complexity: A trade-off in visual question answering models

Moshiur Farazi,Salman Khan,Nick Barnes

doi:10.1016/j.patcog.2021.108106

Abstract

Visual Question Answering (VQA) has emerged as a Visual Turing Test to validate the reasoning ability of AI agents. The pivot to existing VQA models is the joint embedding that is learned by combining the visual features from an image and the semantic features from a given question. Consequently, a large body of literature has focused on developing complex joint embedding strategies coupled with visual attention mechanisms to effectively capture the interplay between these two modalities. However, modelling the visual and semantic features in a high dimensional (joint embedding) space is computationally expensive, and more complex models often result in trivial improvements in the VQA accuracy. In this work, we systematically study the trade-off between the model complexity and the performance on the VQA task. VQA models have a diverse architecture comprising of pre-processing, feature extraction, multimodal fusion, attention and final classification stages. We specifically focus on the effect of “multi-modal fusion” in VQA models that is typically the most expensive step in a VQA pipeline. Our thorough experimental evaluation leads us to three proposals, one optimized for minimal complexity, one for balanced complexity-accuracy and the last one for state-of-the-art VQA performance.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Accuracy vs. complexity: A trade-off in visual question answering models

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition

Lead the way for us

Journal: Pattern Recognition	Publication Date: Jun 12, 2021
Citations: 15

Similar Papers

From known to the unknown: Transferring knowledge to answer questions about novel visual and semantic concepts
Moshiur R Farazi ... Nick Barnes
Image and Vision Computing | VOL. 103
Moshiur R Farazi, et. al.Moshiur R Farazi ... Nick Barnes
04 Aug 2020
Image and Vision Computing | VOL. 103

Neural Networks for Detecting Irrelevant Questions During Visual Question Answering
Mengdi Li ... Cornelius Weber
-
Mengdi Li, et. al.Mengdi Li ... Cornelius Weber
01 Jan 2020
01 Jan 2020

Feature Enhancement in Attention for Visual Question Answering
Yuetan Lin ... Zhangyang Pang
-
Yuetan Lin, et. al.Yuetan Lin ... Zhangyang Pang
01 Jul 2018
01 Jul 2018

Improving Automatic VQA Evaluation Using Large Language Models
Oscar Mañas ... Aishwarya Agrawal
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 38
Oscar Mañas, et. al.Oscar Mañas ... Aishwarya Agrawal
24 Mar 2024
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 38

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Accuracy vs. complexity: A trade-off in visual question answering models

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition