Joint Scence Network and Attention-Guided for Image Captioning

Dongming Zhou,Canlong Zhang,Yanping Tang,Jing Yang

doi:10.1109/icdm51629.2021.00201

Abstract

Image captioning is an interesting and challenging task. The previously established image captioning approach is based mainly on the encoder-decoder architecture, but it suffers from problems such as inaccurate captioning information, and the generated captioning sentences are not sufficiently rich. This paper proposes a novel image captioning model that is based on a self-attention network and a scene graph relationship network. First, an improved self-attention network is added to the extraction of visual features to evaluate the effectiveness of image global information for image generation. Then, we design a visual intensity parameter to coordinate the strategies of visual features and language model for word generation. Finally, a graph convolutional network is designed to extract the relationships from the scene information to render the generated caption more exciting and to increase the accuracy of the fine-grained captioning. We demonstrated the satisfactory performance of the model on the MS-COCO and Flickr 30K datasets. The experimental results demonstrate that the proposed model realizes state-of-the-art performance.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Joint Scence Network and Attention-Guided for Image Captioning

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

A Comparison between Vgg16 and Xception Models used as Encoders for Image Captioning
Asrar Almogbil ... Amjad Alghamdi
-
Asrar Almogbil, et. al.Asrar Almogbil ... Amjad Alghamdi
30 Jul 2022
30 Jul 2022

Normalized and Geometry-Aware Self-Attention Network for Image Captioning
Longteng Guo ... Peng Yao
-
Longteng Guo, et. al.Longteng Guo ... Peng Yao
01 Jun 2020
01 Jun 2020

Multi-Level Policy and Reward-Based Deep Reinforcement Learning Framework for Image Captioning
Ning Xu ... Hanwang Zhang
IEEE Transactions on Multimedia | VOL. 22
Ning Xu, et. al.Ning Xu ... Hanwang Zhang
26 Sep 2019
IEEE Transactions on Multimedia | VOL. 22

TSIC-CLIP: Traffic Scene Image Captioning Model Based on Clip
Hao Zhang ... Xuewei Li
Information Technology and Control | VOL. 53
Hao Zhang, et. al.Hao Zhang ... Xuewei Li
22 Mar 2024
Information Technology and Control | VOL. 53

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Joint Scence Network and Attention-Guided for Image Captioning

Abstract

Talk to us

Similar Papers