Exploring deep learning approaches for video captioning: A comprehensive review

Adel Jalal Yousif,Mohammed H Al-Jammas

doi:10.1016/j.prime.2023.100372

Abstract

While humans can easily describe visual data at varying levels of detail, the same task presents a significant challenge for machines. This challenge becomes even more complex when dealing with video data. The process of understanding a video and generating descriptive text for it is known as video captioning. Video captioning requires not only understanding the visual content but also producing human-like descriptions that accurately capture its semantics. Achieving this level of understanding requires the collaborative efforts of both the computer vision and natural language processing research communities. The captions produced through video captioning serve as valuable resources that can be further leveraged for various applications such as video search, accessibility for visually impaired people, and human-robot interaction. Deep learning strategies have emerged as powerful tools in addressing the complexities of video captioning. By leveraging large scale annotated video caption datasets and sophisticated neural network architectures, deep learning approaches have made significant advances in this challenging task. In the existing literature, numerous techniques, benchmark datasets, and evaluation metrics have been developed, emphasizing the necessity for a comprehensive examination to concentrate research efforts in this rapidly evolving field. This paper provides a survey of deep learning based methods for video captioning, highlighting their key components, challenges, and recent advancements.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: e-Prime - Advances in Electrical Engineering, Electronics and Energy	Publication Date: Nov 22, 2023
Citations: 4	License type: cc-by

R Discovery Prime

R Discovery Prime

Exploring deep learning approaches for video captioning: A comprehensive review

Abstract

Talk to us

Similar Papers

More From: e-Prime - Advances in Electrical Engineering, Electronics and Energy

Lead the way for us

Similar Papers

Video Description
Nayyer Aafaq ... Wei Liu
ACM Computing Surveys | VOL. 52
Nayyer Aafaq, et. al.Nayyer Aafaq ... Wei Liu
16 Oct 2019
ACM Computing Surveys | VOL. 52

A partially flipped physiology classroom improves the deep learning approach of medical students.
Ziqi Liu ... Ziqiang Luo
Advances in physiology education | VOL. 48
Ziqi Liu, et. al.Ziqi Liu ... Ziqiang Luo
11 Apr 2024
Advances in physiology education | VOL. 48

Review Of Video Captioning Methods
Dewarthi Mahajan ... Madhuri Tayal
International Journal of Next-Generation Computing | VOL. -
Dewarthi Mahajan, et. al.Dewarthi Mahajan ... Madhuri Tayal
26 Nov 2021
International Journal of Next-Generation Computing | VOL. -

The Effect of Engagement and Perceived Course Value on Deep and Surface Learning Strategies
Kevin Floyd ... Susan Harrington
-
Kevin Floyd, et. al.Kevin Floyd ... Susan Harrington
01 Jan 2009
01 Jan 2009

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Exploring deep learning approaches for video captioning: A comprehensive review

Abstract

Talk to us

Similar Papers

More From: e-Prime - Advances in Electrical Engineering, Electronics and Energy