UNISON: Unpaired Cross-Lingual Image Captioning

Jiahui Gao,Jiuxiang Gu,Philip L H Yu,Yi Zhou,Shafiq Joty

doi:10.1609/aaai.v36i10.21310

Abstract

Image captioning has emerged as an interesting research field in recent years due to its broad application scenarios. The traditional paradigm of image captioning relies on paired image-caption datasets to train the model in a supervised manner. However, creating such paired datasets for every target language is prohibitively expensive, which hinders the extensibility of captioning technology and deprives a large part of the world population of its benefit. In this work, we present a novel unpaired cross-lingual method to generate image captions without relying on any caption corpus in the source or the target language. Specifically, our method consists of two phases: (1) a cross-lingual auto-encoding process, which utilizing a sentence parallel (bitext) corpus to learn the mapping from the source to the target language in the scene graph encoding space and decode sentences in the target language, and (2) a cross-modal unsupervised feature mapping, which seeks to map the encoded scene graph features from image modality to language modality. We verify the effectiveness of our proposed method on the Chinese image caption generation task. The comparisons against several existing methods demonstrate the effectiveness of our approach.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

UNISON: Unpaired Cross-Lingual Image Captioning

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Jun 28, 2022
Citations: 3

Similar Papers

Topic Scene Graph Generation by Attention Distillation from Caption
Wenbin Wang ... Xilin Chen
-
Wenbin Wang, et. al.Wenbin Wang ... Xilin Chen
01 Oct 2021
01 Oct 2021

Visual-linguistic-stylistic Triple Reward for Cross-lingual Image Captioning
Jing Zhang ... Meng Wang
ACM Transactions on Multimedia Computing, Communications, and Applications | VOL. 20
Jing Zhang, et. al.Jing Zhang ... Meng Wang
11 Jan 2024
ACM Transactions on Multimedia Computing, Communications, and Applications | VOL. 20

Synthesis of Vision and Language: Multifaceted Image Captioning Application
Arpit Gupta ... Ishita Kohli
INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT | VOL. 07
Arpit Gupta, et. al.Arpit Gupta ... Ishita Kohli
23 Dec 2023
INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT | VOL. 07

In Defense of Scene Graphs for Image Captioning
Kien Nguyen ... Tanaya Guha
-
Kien Nguyen, et. al.Kien Nguyen ... Tanaya Guha
01 Oct 2021
01 Oct 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

UNISON: Unpaired Cross-Lingual Image Captioning

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence