Coupled feature mapping and correlation mining for cross-media retrieval

Mengdi Fan Mengdi Fan,Ronggang Wang Ronggang Wang,Wenmin Wang

doi:10.1109/icmew.2016.7574754

Abstract

Cross-media retrieval aims to integrate and analyze the features of various modalities (e.g., text, image and video) to mine their potential semantic information. In this paper, we propose a novel cross-media retrieval framework, which performs coupled feature mapping and correlation mining successively. Our method first learns two projection matrices to map the multimodal features into a common category space, in which homo- and hetero-correlation techniques can be applied easily. Homo-correlation focuses on the semantic category information within the same media type, while hetero-correlation focuses on the semantic category information be-tween different media types. The two could complement and reinforce each other. Experiments on two different datasets, Wikipedia dataset and Pascal Voc dataset, demonstrate that the proposed framework gives promising results compared to the related state-of-the-art approaches.

Full Text