Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion

Jinpeng Wang,Xing Sun,Xiaowei Guo,Yuting Gao,Jianguo Hu,Rongrong Ji,Ke Li,Xinyang Jiang

doi:10.1609/aaai.v35i11.17215

Abstract

One significant factor we expect the video representation learning to capture, especially in contrast with the image representation learning, is the object motion. However, we found that in the current mainstream video datasets, some action categories are highly related with the scene where the action happens, making the model tend to degrade to a solution where only the scene information is encoded. For example, a trained model may predict a video as playing football simply because it sees the field, neglecting that the subject is dancing as a cheerleader on the field. This is against our original intention towards the video representation learning and may bring scene bias on a different dataset that can not be ignored. In order to tackle this problem, we propose to decouple the scene and the motion (DSM) with two simple operations, so that the model attention towards the motion information is better paid. Specifically, we construct a positive clip and a negative clip for each video. Compared to the original video, the positive/negative is motion-untouched/broken but scene-broken/untouched by Spatial Local Disturbance and Temporal Local Disturbance. Our objective is to pull the positive closer while pushing the negative farther to the original clip in the latent space. In this way, the impact of the scene is weakened while the temporal sensitivity of the network is further enhanced. We conduct experiments on two tasks with various backbones and different pre-training datasets, and find that our method surpass the SOTA methods with a remarkable 8.1% and 8.8% improvement towards action recognition task on the UCF101 and HMDB51 datasets respectively using the same backbone.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: May 18, 2021
Citations: 25

Similar Papers

VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
Tian Pan ... Wei Liu
-
Tian Pan, et. al.Tian Pan ... Wei Liu
01 Jun 2021
01 Jun 2021

Learning Unsupervised Visual Representations using 3D Convolutional Autoencoder with Temporal Contrastive Modeling for Video Retrieval
Vidit Kumar ... Vikas Tripathi
International Journal of Mathematical, Engineering and Management Sciences | VOL. 7
Vidit Kumar, et. al.Vidit Kumar ... Vikas Tripathi
14 Mar 2022
International Journal of Mathematical, Engineering and Management Sciences | VOL. 7

SSAN: Separable Self-Attention Network for Video Representation Learning
Xudong Guo ... Yan Lu
-
Xudong Guo, et. al.Xudong Guo ... Yan Lu
01 Jun 2021
01 Jun 2021

Motion-aware Contrastive Video Representation Learning via Foreground-background Merging
Shuangrui Ding ... Hongkai Xiong
-
Shuangrui Ding, et. al.Shuangrui Ding ... Hongkai Xiong
01 Jun 2022
01 Jun 2022

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence