VISCOUNTH: A Large-scale Multilingual Visual Question Answering Dataset for Cultural Heritage

Federico Becattini,Luana Bulla,Misael Mongiovì,Ludovica Marinucci,Alberto Del Bimbo,Valentina Presutti,Pietro Bongini

doi:10.1145/3590773

Abstract

Visual question answering has recently been settled as a fundamental multi-modal reasoning task of artificial intelligence that allows users to get information about visual content by asking questions in natural language. In the cultural heritage domain, this task can contribute to assisting visitors in museums and cultural sites, thus increasing engagement. However, the development of visual question answering models for cultural heritage is prevented by the lack of suitable large-scale datasets. To meet this demand, we built a large-scale heterogeneous and multilingual (Italian and English) dataset for cultural heritage that comprises approximately 500K Italian cultural assets and 6.5M question-answer pairs. We propose a novel formulation of the task that requires reasoning over both the visual content and an associated natural language description, and present baselines for this task. Results show that the current state of the art is reasonably effective but still far from satisfactory; therefore, further research in this area is recommended. Nonetheless, we also present a holistic baseline to address visual and contextual questions and foster future research on the topic.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: ACM Transactions on Multimedia Computing, Communications, and Applications	Publication Date: Jul 12, 2023
Citations: 3	License type: cc-by-sa

R Discovery Prime

R Discovery Prime

VISCOUNTH: A Large-scale Multilingual Visual Question Answering Dataset for Cultural Heritage

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Multimedia Computing, Communications, and Applications

Lead the way for us

Similar Papers

Visual Question Answering as Reading Comprehension
Hui Li ... Peng Wang
-
Hui Li, et. al.Hui Li ... Peng Wang
01 Jun 2019
01 Jun 2019

Visual Question Answering for Cultural Heritage
Pietro Bongini ... Alberto Del Bimbo
IOP Conference Series: Materials Science and Engineering | VOL. 949
Pietro Bongini, et. al.Pietro Bongini ... Alberto Del Bimbo
01 Nov 2020
IOP Conference Series: Materials Science and Engineering | VOL. 949

Prior Visual Relationship Reasoning For Visual Question Answering
Zhuoqian Yang ... Jing Yu
-
Zhuoqian Yang, et. al.Zhuoqian Yang ... Jing Yu
01 Oct 2020
01 Oct 2020

Estimating Viewed Images with Natural Language Question Answering from fMRI Data
Saya Takada ... Takahiro Ogawa
-
Saya Takada, et. al.Saya Takada ... Takahiro Ogawa
01 Mar 2020
01 Mar 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

VISCOUNTH: A Large-scale Multilingual Visual Question Answering Dataset for Cultural Heritage

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Multimedia Computing, Communications, and Applications