Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration

Ngoc Son Nguyen,Van Son Nguyen,Tung Le

doi:10.1016/j.compeleceng.2024.109474

Abstract

Visual Question Answering (VQA) has recently emerged as a potential research domain, captivating the interest of many in the field of artificial intelligence and computer vision. Despite the prevalence of approaches in English, there is a notable lack of systems specifically developed for certain languages, particularly Vietnamese. This study aims to bridge this gap by conducting comprehensive experiments on the Vietnamese Visual Question Answering (ViVQA) dataset, demonstrating the effectiveness of our proposed model. In response to community interest, we have developed a model that enhances image representation capabilities, thereby improving overall performance in the ViVQA system. Therefore, we propose AViVQA-TranConI (Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration). AViVQA-TranConI integrates the Bootstrapping Language-Image Pre-training with frozen unimodal models (BLIP-2) and the convolutional neural network EfficientNet to extract and process both local and global features from images. This integration leverages the strengths of transformer-based architectures for capturing comprehensive contextual information and convolutional networks for detailed local features. By freezing the parameters of these pre-trained models, we significantly reduce the computational cost and training time, while maintaining high performance. This approach significantly improves image representation and enhances the performance of existing VQA systems. We then leverage a multi-modal fusion module based on a general-purpose multi-modal foundation model (BEiT-3) to fuse the information between visual and textual features. Our experimental findings demonstrate that AViVQA-TranConI surpasses competing baselines, achieving promising performance. This is particularly evident in its accuracy of 71.04% on the test set of the ViVQA dataset, marking a significant advancement in our research area. The code is available at https://github.com/nngocson2002/ViVQA.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration

Abstract

Talk to us

Similar Papers

More From: Computers and Electrical Engineering

Lead the way for us

Similar Papers

Visual Question Answering Based on Question Attention Model
Jianing Zhang ... Huajie Zhang
Journal of Physics: Conference Series | VOL. 1624
Jianing Zhang, et. al.Jianing Zhang ... Huajie Zhang
01 Oct 2020
Journal of Physics: Conference Series | VOL. 1624

Joint Coding of Local and Global Deep Features in Videos for Visual Search.
Lin Ding ... Tiejun Huang
IEEE Transactions on Image Processing | VOL. 29
Lin Ding, et. al.Lin Ding ... Tiejun Huang
01 Jan 2020
IEEE Transactions on Image Processing | VOL. 29

Patch Embedding as Local Features: Unifying Deep Local and Global Features via Vision Transformer for Image Retrieval
Lam Phan ... Harikrishna Warrier
-
Lam Phan, et. al.Lam Phan ... Harikrishna Warrier
01 Jan 2023
01 Jan 2023

CNN classification based on global and local features
Yufeng Zheng ... Matthias F Carlsohn
-
Yufeng Zheng, et. al.Yufeng Zheng ... Matthias F Carlsohn
14 May 2019
14 May 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration

Abstract

Talk to us

Similar Papers

More From: Computers and Electrical Engineering