Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

Vincent Cartillier,Stefan Lee,Neha Jain,Zhile Ren,Irfan Essa,Dhruv Batra

doi:10.1609/aaai.v35i2.16180

Abstract

We study the task of semantic mapping – specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (‘what is where?’) from egocentric observations of an RGB-D camera with known pose (via localization sensors). Importantly, our goal is to build neural episodic memories and spatio-semantic representations of 3D spaces that enable the agent to easily learn subsequent tasks in the same space – navigating to objects seen during the tour (‘Find chair’) or answering questions about the space (‘How many chairs did you see in the house?’). Towards this goal, we present Semantic MapNet (SMNet), which consists of: (1) an Egocentric Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length×width×feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01−16.81% (absolute) on mean-IoU and 3.81−19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the spatio-semantic allocentric representations build by SMNet for the task of ObjectNav and Embodied Question Answering. Project page: https://vincentcartillier.github.io/smnet.html.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: May 18, 2021
Citations: 33

Similar Papers

Spatial semantic hybrid map building and application of mobile service robot
Hao Wu ... Peng Duan
Robotics and Autonomous Systems | VOL. 62
Hao Wu, et. al.Hao Wu ... Peng Duan
23 Jan 2013
Robotics and Autonomous Systems | VOL. 62

Modeling the Development of Lexicon with DevLex: A Self-Organizing Neural Network Model of Lexical Acquisition
Ping Li ... Igor Farkaš
-
Ping Li, et. al.Ping Li ... Igor Farkaš
24 Apr 2019
24 Apr 2019

Principal semantic components of language and the measurement of meaning.
Giorgio A Ascoli ... Alexei V Samsonovic
PLoS ONE | VOL. 5
Giorgio A Ascoli, et. al.Giorgio A Ascoli ... Alexei V Samsonovic
11 Jun 2010
PLoS ONE | VOL. 5

SEMANTIC MAPPING FROM NATURAL LANGUAGE QUESTIONS TO OWL QUERIES
Mingxia Gao ... Jiming Liu
Computational Intelligence | VOL. 27
Mingxia Gao, et. al.Mingxia Gao ... Jiming Liu
01 May 2011
Computational Intelligence | VOL. 27

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence