Depth and Video Segmentation Based Visual Attention for Embodied Question Answering.

Haonan Luo,Zhenmin Tang,Fayao Liu,Zichuan Liu,Guosheng Lin,Yazhou Yao

doi:10.1109/tpami.2021.3139957

Abstract

Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the users questions by exploring the real-world environment. It has attracted increasing research interests due to its broad applications in personal assistants and in-home robots. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of fine-level semantic information, stability to the ambiguity, and 3D spatial information of the virtual environment. To tackle these problems, we propose a depth and segmentation based visual attention mechanism for Embodied Question Answering. Firstly, we extract local semantic features by introducing a novel high-speed video segmentation framework. Then guided by the extracted semantic features, a depth and segmentation based visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is designed to guide the navigators training process without much additional computational cost. The ablation experiments show that our method effectively boosts the performance of the VQA module and navigation module, leading to 4.9% and 5.6% overall improvement in EQA accuracy on House3D and Matterport3D datasets respectively.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Depth and Video Segmentation Based Visual Attention for Embodied Question Answering.

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Pattern Analysis and Machine Intelligence

Lead the way for us

Journal: IEEE Transactions on Pattern Analysis and Machine Intelligence	Publication Date: Jun 1, 2023
Citations: 6

Similar Papers

SegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering
Haonan Luo ... Zichuan Liu
-
Haonan Luo, et. al.Haonan Luo ... Zichuan Liu
01 Oct 2019
01 Oct 2019

Accuracy vs. complexity: A trade-off in visual question answering models
Moshiur Farazi ... Nick Barnes
Pattern Recognition | VOL. 120
Moshiur Farazi, et. al.Moshiur Farazi ... Nick Barnes
12 Jun 2021
Pattern Recognition | VOL. 120

Robust-EQA: Robust Learning for Embodied Question Answering With Noisy Labels.
Haonan Luo ... Hengtao Shen
IEEE transactions on neural networks and learning systems | VOL. 35
Haonan Luo, et. al.Haonan Luo ... Hengtao Shen
01 Sep 2024
IEEE transactions on neural networks and learning systems | VOL. 35

Co-Attending Free-Form Regions and Detections With Multi-Modal Multiplicative Feature Embedding for Visual Question Answering
Pan Lu ... Jianyong Wang
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 32
Pan Lu, et. al.Pan Lu ... Jianyong Wang
27 Apr 2018
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 32

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Depth and Video Segmentation Based Visual Attention for Embodied Question Answering.

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Pattern Analysis and Machine Intelligence