Multimodal Learning with Triplet Ranking Loss for Visual Semantic Embedding Learning

Zhanbo Yang,Li Li,Jun He,Li Liu,Jun Liao,Zixi Wei

doi:10.1007/978-3-030-29551-6_67

Abstract

Semantic embedding learning for image and text has been well studied in recent years. In this paper, we present a simple while effective dual-encoder (image encoder and text encoder) framework to unify image and text into a common embedding space. Inspired by deep metric learning, we utilize triplet ranking loss to minimize the gap between the two embedding spaces. We train and test our proposed framework on Flickr8k, Flickr30k and MS-COCO datasets respectively, and evaluate the framework on the Corel1k benchmark dataset as an application. Using VGG-19 for image encoder, GRU for text encoder and triplet ranking loss, we gained obvious improvement versus baseline model on image annotation and image search tasks. Additionally, we explore the vector generated by our image encoder and the one by word embedding of plain word for some arithmetic operations. The above experiments demonstrate the effectiveness of our proposed learning framework.

Full Text