Hybrid-attention and frame difference enhanced network for micro-video venue recognition

Bing Wang,Xianglin Huang,Gang Cao,Lifang Yang,Zhulin Tao,Xiaolong Wei

doi:10.3233/jifs-213191

Abstract

Many micro-video related applications, such as personalized location recommendation and micro-video verification, can be benefited greatly from the venue information. Most existing works focus on integrating the information from multi-modal for exact venue category recognition. It is important to make full use of the information from different modalities. However, the performance may be limited by the lacked acoustic modality or textual descriptions in uploaded micro-videos. Therefore, in this paper visual modality is explored as the only modality according to its rich and indispensable semantic information. To this end, a hybrid-attention and frame difference enhanced network (HAFDN) is proposed to generate the comprehensive venue representation. Such network mainly contains two parallel branches: content and motion branches. Specifically, in the content branch, a domain-adaptive CNN model combined with temporal shift module (TSM) is employed to extract discriminative visual features. Then, a novel hybrid attention module (HAM) is introduced to enhance extracted features via three attention mechanisms. In HAM, channel attention, local and global spatial attention mechanisms are used to capture salient visual information from different views. In addition, convolutional Long Short-Term Memory (convLSTM) is enforced after HAM to better encode the long spatial-temporal dependency. A difference-enhanced module parallel with HAM is devised to learn the content variations among adjacent frames, which is usually ignored in prior works. Moreover, in the motion branch, 3D-CNNs and LSTM are used to capture movement variation as a supplement of content branch in a different form. Finally, the features from two branches are fused to generate robust video-level representations for predicting venue categories. Extensive experimental results on public datasets verify the effectiveness of the proposed micro-video venue recognition scheme. The source code is available at https://github.com/hs8945/HAFDN.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Hybrid-attention and frame difference enhanced network for micro-video venue recognition

Abstract

Talk to us

Similar Papers

More From: Journal of Intelligent & Fuzzy Systems

Lead the way for us

Journal: Journal of Intelligent & Fuzzy Systems	Publication Date: Jul 21, 2022
Citations: 2

Similar Papers

EFFNet-CA: An Efficient Driver Distraction Detection Based on Multiscale Features Extractions and Channel Attention Mechanism.
Taimoor Khan ... Gyuho Choi
Sensors (Basel, Switzerland) | VOL. 23
Taimoor Khan, et. al.Taimoor Khan ... Gyuho Choi
08 Apr 2023
Sensors (Basel, Switzerland) | VOL. 23

Spatial-Channel Transformer for Scene Recognition
Seunghyun Baik ... Euntai Kim
-
Seunghyun Baik, et. al.Seunghyun Baik ... Euntai Kim
18 Jul 2022
18 Jul 2022

Lightweight Image Super-resolution with Local Attention Enhancement
Yunchu Yang ... Xiumei Wang
-
Yunchu Yang, et. al.Yunchu Yang ... Xiumei Wang
01 Jan 2020
01 Jan 2020

FcaNet: Frequency Channel Attention Networks
Zequn Qin ... Fei Wu
-
Zequn Qin, et. al.Zequn Qin ... Fei Wu
01 Oct 2021
01 Oct 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Hybrid-attention and frame difference enhanced network for micro-video venue recognition

Abstract

Talk to us

Similar Papers

More From: Journal of Intelligent &amp; Fuzzy Systems

More From: Journal of Intelligent & Fuzzy Systems