Topic-aware video summarization using multimodal transformer

Yubo Zhu,Wentian Zhao,Rui Hua,Xinxiao Wu

doi:10.1016/j.patcog.2023.109578

Abstract

Video summarization aims to generate a short and compact summary to represent the original video. Existing methods mainly focus on how to extract a general objective synopsis that precisely summaries the video content. However, in real scenarios, a video usually contains rich content with multiple topics and people may cast diverse interests on the visual contents even for the same video. In this paper, we propose a novel topic-aware video summarization task that generates multiple video summaries with different topics. To support the study of this new task, we first build a video benchmark dataset by collecting videos from various types of movies and annotate them with topic labels and frame-level importance scores. Then we propose a multimodal Transformer model for the topic-aware video summarization, which simultaneously predicts topic labels and generates topic-related summaries by adaptively fusing multimodal features extracted from the video. Experimental results show the effectiveness of our method.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Topic-aware video summarization using multimodal transformer

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition

Lead the way for us

Journal: Pattern Recognition	Publication Date: Mar 30, 2023
Citations: 9

Similar Papers

A comprehensive survey and mathematical insights towards video summarization
Pulkit Narwal ... Komal Kumar Bhatia
Journal of Visual Communication and Image Representation | VOL. 89
Pulkit Narwal, et. al.Pulkit Narwal ... Komal Kumar Bhatia
17 Oct 2022
Journal of Visual Communication and Image Representation | VOL. 89

Learning user interest with improved triplet deep ranking and web-image priors for topic-related video summarization
Mengjuan Fei ... Weijie Mao
Expert Systems with Applications | VOL. 166
Mengjuan Fei, et. al.Mengjuan Fei ... Weijie Mao
29 Sep 2020
Expert Systems with Applications | VOL. 166

ILS-SUMM: Iterated Local Search for Unsupervised Video Summarization
Yair Shemer ... Daniel Rotman
-
Yair Shemer, et. al.Yair Shemer ... Daniel Rotman
10 Jan 2021
10 Jan 2021

Unsupervised Multi-Topic Labeling for Spoken Utterances
Sebastian Weigelt ... Jan Keim
-
Sebastian Weigelt, et. al.Sebastian Weigelt ... Jan Keim
01 Sep 2019
01 Sep 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Topic-aware video summarization using multimodal transformer

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition