Spatiotemporal Pyramid Network for Video Action Recognition

Yunbo Wang,Mingsheng Long,Jianmin Wang,Philip S Yu

doi:10.1109/cvpr.2017.226

Abstract

Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact bilinear operator into video analysis tasks. This operator enables efficient training of bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Spatiotemporal Pyramid Network for Video Action Recognition

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Improved two-stream model for human action recognition
Yuxuan Zhao ... Kamran Siddique
EURASIP Journal on Image and Video Processing | VOL. 2020
Yuxuan Zhao, et. al.Yuxuan Zhao ... Kamran Siddique
17 Jun 2020
EURASIP Journal on Image and Video Processing | VOL. 2020

Spatial-temporal interaction learning based two-stream network for action recognition
Tianyu Liu ... Ping Jiang
Information Sciences | VOL. 606
Tianyu Liu, et. al.Tianyu Liu ... Ping Jiang
28 May 2022
Information Sciences | VOL. 606

Multi-scale Spatiotemporal Information Fusion Network for Video Action Recognition
Yutong Cai ... Ming-Ming Cheng
-
Yutong Cai, et. al.Yutong Cai ... Ming-Ming Cheng
01 Dec 2018
01 Dec 2018

R(2+1)D-based Two-stream CNN for Human Activities Recognition in Videos
Min Huang ... Yi Han
-
Min Huang, et. al.Min Huang ... Yi Han
26 Jul 2021
26 Jul 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Spatiotemporal Pyramid Network for Video Action Recognition

Abstract

Talk to us

Similar Papers