Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption

Wei Zhang,Yue Ying,Pan Lu,Hongyuan Zha

doi:10.1609/aaai.v34i05.6503

Abstract

Personalized image caption, a natural extension of the standard image caption task, requires to generate brief image descriptions tailored for users' writing style and traits, and is more practical to meet users' real demands. Only a few recent studies shed light on this crucial task and learn static user representations to capture their long-term literal-preference. However, it is insufficient to achieve satisfactory performance due to the intrinsic existence of not only long-term user literal-preference, but also short-term literal-preference which is associated with users' recent states. To bridge this gap, we develop a novel multimodal hierarchical transformer network (MHTN) for personalized image caption in this paper. It learns short-term user literal-preference based on users' recent captions through a short-term user encoder at the low level. And at the high level, the multimodal encoder integrates target image representations with short-term literal-preference, as well as long-term literal-preference learned from user IDs. These two encoders enjoy the advantages of the powerful transformer networks. Extensive experiments on two real datasets show the effectiveness of considering two types of user literal-preference simultaneously and better performance over the state-of-the-art models.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Apr 3, 2020
Citations: 17

Similar Papers

Multimodal Transformer With Multi-View Visual Representation for Image Captioning
Jun Yu ... Qingming Huang
IEEE Transactions on Circuits and Systems for Video Technology | VOL. 30
Jun Yu, et. al.Jun Yu ... Qingming Huang
25 Oct 2019
IEEE Transactions on Circuits and Systems for Video Technology | VOL. 30

Research on Image Caption Based on Multiple Word Embedding Representations
Zhen-Xian Lin ... Xiao-Bao Yang
-
Zhen-Xian Lin, et. al.Zhen-Xian Lin ... Xiao-Bao Yang
01 Mar 2021
01 Mar 2021

Maternal and genetic influences on production and reproduction traits in pigs
H.A.M Van Der Steen
Netherlands Journal of Agricultural Science | VOL. 32
H.A.M Van Der SteenH.A.M Van Der Steen
01 Feb 1984
Netherlands Journal of Agricultural Science | VOL. 32

Two Distinct Visual Motion Mechanisms for Smooth Pursuit: Evidence from Individual Differences
Jeremy B Wilmer ... Ken Nakayama
Neuron | VOL. 54
Jeremy B Wilmer, et. al.Jeremy B Wilmer ... Ken Nakayama
01 Jun 2007
Neuron | VOL. 54

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence