Learning Spatiotemporal-Selected Representations in Videos for Action Recognition

Jiachao Zhang,Liangbao Jiao,Ying Tong

doi:10.1142/s0218126623502031

Abstract

Action recognition is a challenging task of modeling both spatial and temporal context. Numerous works focus on architectures modality and successfully make worthy progress on this task. While due to the redundancy in time and the limit of computation resources, several works focus on the efficiency study like frame sampling, some for untrimmed videos, and some for trimmed videos. With the intent of improving the effectiveness of action recognition, we propose a novel Computational Spatiotemporal Selector (CSS) to refine and reinforce the key frames with discriminative information in video. Specifically, CSS includes two modules: Temporal Adaptive Sampling (TAS) module and Spatial Frame Resolution (SFR) module. The former can refine the key frames in the temporal space for capturing the key motion information, while the latter can further zoom out some refined frames in the spatial space for eliminating the discrimination-irrelevant structural information. The proposed CSS is flexible to be embedded into most representative action recognition models. Experiments on two challenging action recognition benchmarks, i.e., ActivityNet1.3 and UCF101, show that the proposed CSS improves the performance over most existing models, not only on trimmed videos but also untrimmed videos.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Learning Spatiotemporal-Selected Representations in Videos for Action Recognition

Abstract

Talk to us

Similar Papers

More From: Journal of Circuits, Systems and Computers

Lead the way for us

Similar Papers

Action recognition and localization with spatial and temporal contexts
Wanru Xu ... Qiang Ji
Neurocomputing | VOL. 333
Wanru Xu, et. al.Wanru Xu ... Qiang Ji
09 Jan 2019
Neurocomputing | VOL. 333

Action Recognition in Untrimmed Videos with Composite Self-attention Two-Stream Framework
Dong Cao ... Lisha Xu
-
Dong Cao, et. al.Dong Cao ... Lisha Xu
01 Jan 2020
01 Jan 2020

Spatio-Temporal Activity Detection and Recognition in Untrimmed Surveillance Videos
Konstantinos Gkountakos ... Ioannis Kompatsiaris
-
Konstantinos Gkountakos, et. al.Konstantinos Gkountakos ... Ioannis Kompatsiaris
24 Aug 2021
24 Aug 2021

Action recognition on continuous video
Y L Chang ... C S Chan
Neural Computing and Applications | VOL. 33
Y L Chang, et. al.Y L Chang ... C S Chan
05 Jun 2020
Neural Computing and Applications | VOL. 33

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Learning Spatiotemporal-Selected Representations in Videos for Action Recognition

Abstract

Talk to us

Similar Papers

More From: Journal of Circuits, Systems and Computers