Violence 4D: Violence detection in surveillance using 4D convolutional neural networks

Mai Magdy,Fahima A Maghraby,Mohamed Waleed Fakhr

doi:10.1049/cvi2.12162

Mai Magdy, Fahima A Maghraby + Show 1 more

Open Access

https://doi.org/10.1049/cvi2.12162

Copy DOI

Abstract

AbstractAs violence has increased around the world, surveillance cameras are everywhere, and they are only going to get more ubiquitous. Due to the massive volume of video footage, automatic activity detection systems must be used to create an online warning in the event of aberrant activity. A deep learning architecture is presented in this study using four‐dimensional video‐level convolution neural networks. The proposed architecture includes residual blocks that are used with three‐Dimensional Convolution Neural Networks 3D (CNNs) to learn long‐term and short‐term spatiotemporal representation from the video as well as record inter‐clip interaction. ResNet50 is used as the backbone for three‐dimensional convolution networks and dense optical flow for the region of interest. The proposed architecture is applied on four benchmarks for violence and non‐violence videos, which are commonly used for violent detection. It obtained test accuracies of 94.67% on RWF2000, 97.29% on Crowd violence, 100% on Movie fight and 100% on the Hockey Fight dataset. These results outperform the previous methods used on RWF2000 datasets.

Full Text