Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Spatiotemporal deep learning framework for predictive behavioral threat detection in surveillance footage

  • Abstract
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Anomaly detection in video surveillance remains a challenging problem due to complex human behaviors, temporal variability, and limited annotated data. This study proposes an optimized spatiotemporal deep learning (DL) framework that integrates a Convolutional Neural Network (CNN) for spatial feature extraction with a Long Short-Term Memory (LSTM) network for temporal dependency modeling. The CNN processes frame-level appearance information, while the LSTM captures sequential motion patterns across video frames, enabling effective representation of anomalous activities. Hyperparameter optimization and regularization strategies are employed to improve convergence stability and generalization performance. The proposed model is evaluated on the DCSASS surveillance dataset and the experimental results demonstrate that the optimized CNN-LSTM framework achieves an accuracy of 98.1%, with consistently high precision, recall, and F1-score across 3-fold, 5-fold, and 10-fold cross-validation settings. Comparative analysis shows that the proposed method outperforms conventional machine learning models and recent deep learning baselines, highlighting its effectiveness and robustness for practical video-based anomaly detection in surveillance environments.

Similar Papers
  • Research Article
  • 10.48084/etasr.10656
Early Anomalus Action Detection in Surveillance Video Using MRCNN-LSTM Classification
  • Aug 2, 2025
  • Engineering, Technology & Applied Science Research
  • D Manju + 7 more

Public space monitoring systems are critical for observing typical human behavior and detecting abnormal activities, especially in high-security environments. With the rise in public space thefts, there is a growing need for intelligent systems capable of detecting suspicious movements early enough to prevent criminal acts. Although Convolutional Neural Networks (CNNs) are widely used in image classification, they are inadequate to differentiate between abnormal and normal behavior and identify criminal activity in its early stage. To overcome these limitations, this study proposes a new hybrid model that combines Mask R-CNN (MRCNN) with Long Short-Term Memory (LSTM) networks for accurate object detection, tracking, and sequential behavior analysis. The main contribution of this study is a multistage anomaly detection pipeline that involves frame conversion, contrast enhancement, background removal, object tracking, and feature extraction. The MRCNN-LSTM framework can extract both spatial and temporal characteristics to allow precise early-stage anomaly detection. Thorough testing on three benchmarking datasets, UCF Crime, Snatch1.0, and CUHK, exhibited excellent performance, with a 93.6% accuracy for the UCF Crime dataset. Performance metrics such as observation ratio and time duration were used to assess the responsiveness and effectiveness of the system in real-time surveillance scenarios. This research advances the field of intelligent surveillance by enabling proactive threat mitigation through the early and precise detection of anomalous behavior.

  • Research Article
  • Cite Count Icon 31
  • 10.1016/j.epsr.2022.109065
Online leakage current classification using convolutional neural network long short-term memory for high voltage insulators on web-based service
  • Mar 1, 2023
  • Electric Power Systems Research
  • Phuong Nguyen Thanh + 1 more

Online leakage current classification using convolutional neural network long short-term memory for high voltage insulators on web-based service

  • Book Chapter
  • Cite Count Icon 1
  • 10.1007/978-3-030-37429-7_21
Global Anomaly Detection Based on a Deep Prediction Neural Network
  • Jan 1, 2019
  • Ang Li + 5 more

Abnormal event detection in public scenes is very important in recent society. In this paper, a method for global anomaly detection in video surveillance is proposed, which is based on a deep prediction neural network. The deep prediction neural network is built on the Convolutional Neural Network (CNN) and a variant of the Recurrent Neural Network (RNN)-Long Short-Term Memory (LSTM). Especially, the feature of a frame is the output of CNN, which is instead of the hand-crafted feature. First, the feature of a short video clip is obtained through CNN. Second, the predicted feature of the next frame can be gained by LSTM. Finally, the prediction error is introduced to detect that a frame is abnormal or not after the feature of the frame is achieved. Experimental results of global abnormal event detection show the effectiveness of our deep prediction neural network. Comparing with state-of-the-art methods, the model we proposed obtains superior detection results.

  • Research Article
  • Cite Count Icon 74
  • 10.1016/j.jhydrol.2023.129401
Improving LSTM hydrological modeling with spatiotemporal deep learning and multi-task learning: A case study of three mountainous areas on the Tibetan Plateau
  • Mar 15, 2023
  • Journal of Hydrology
  • Bu Li + 6 more

Improving LSTM hydrological modeling with spatiotemporal deep learning and multi-task learning: A case study of three mountainous areas on the Tibetan Plateau

  • Research Article
  • Cite Count Icon 3
  • 10.22214/ijraset.2024.64181
A Comprehensive Comparative Analysis of Deep Learning Architectures for Suspicious Activity Detection in Video Surveillance
  • Sep 30, 2024
  • International Journal for Research in Applied Science and Engineering Technology
  • Kashish D Thanki + 2 more

Abstract: In the landscape of modern security infrastructure, video surveillance has evolved into a ubiquitous and indispensable tool for ensuring public safety and safeguarding critical assets. This research paper delves into an extensive examination of various deep learning architectures for the purpose of detecting suspicious activities in video surveillance. The investigation encompasses Convolutional Long Short-Term Memory (ConvLSTM), Convolutional Neural Network (CNN) combined with Long Short-Term Memory (LSTM), ConvLSTM, Bidirectional Long Short-Term Memory (BiLSTM), and diverse combinations thereof. Each model is rigorously trained and tested on a carefully curated dataset designed to encapsulate a spectrum of normal and suspicious activities like shooting and fighting. The study aspires to identify the most efficacious model for enhancing the accuracy of suspicious activity detection in complex and dynamic environments. Metrics such as accuracy, precision, recall, and F1 score will be rigorously assessed to ascertain the model’s comparative performances

  • Research Article
  • 10.70968/ijeaca.v2i1.d1004
Anomaly Detection in Videos Using Deep Learning
  • Jun 28, 2025
  • International Journal of Electronics and Computer Applications
  • Ashwini Patil

In recent decades, surveillance cameras have been widely deployed in various locations for security and monitoring. The video data analysis captured by these cameras plays a crucial role in event prediction, real-time tracking, and goal-driven applications such as anomaly and intrusion detection. With advancements in Artificial Intelligence (AI), deep learning-based approaches, particularly Convolutional Neural Networks (CNNs), have significantly improved anomaly detection accuracy. This study proposes a novel deep-learning framework for detecting anomalies in video surveillance, leveraging CNNs for spatial feature extraction. In recent decades, surveillance cameras have been widely deployed in various locations for security and monitoring. The video data analysis captured by these cameras plays a crucial role in event prediction, real-time tracking, and goal-driven applications such as anomaly and intrusion detection. With advancements in Artificial Intelligence (AI), deep learning-based approaches, particularly Convolutional Neural Networks (CNNs), have significantly improved anomaly detection accuracy. This study proposes a novel deep-learning framework for detecting anomalies in video surveillance, leveraging CNNs for spatial feature extraction. The approach is inspired by previous studies, such as "Anomaly Detection in Surveillance Videos Using Deep Learning" 1, which demonstrated high efficiency in recognizing abnormal patterns. The UCSD dataset has been used to assess the suggested approach, showing improved accuracy in anomaly detection compared to existing methods. The findings highlight the potential of deep learning in enhancing automated surveillance systems, contributing to intelligent security monitoring and public safety. 1

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 18
  • 10.3390/a17070286
Enhancing Video Anomaly Detection Using a Transformer Spatiotemporal Attention Unsupervised Framework for Large Datasets
  • Jul 1, 2024
  • Algorithms
  • Mohamed H Habeb + 2 more

This work introduces an unsupervised framework for video anomaly detection, leveraging a hybrid deep learning model that combines a vision transformer (ViT) with a convolutional spatiotemporal relationship (STR) attention block. The proposed model addresses the challenges of anomaly detection in video surveillance by capturing both local and global relationships within video frames, a task that traditional convolutional neural networks (CNNs) often struggle with due to their localized field of view. We have utilized a pre-trained ViT as an encoder for feature extraction, which is then processed by the STR attention block to enhance the detection of spatiotemporal relationships among objects in videos. The novelty of this work is utilizing the ViT with the STR attention to detect video anomalies effectively in large and heterogeneous datasets, an important thing given the diverse environments and scenarios encountered in real-world surveillance. The framework was evaluated on three benchmark datasets, i.e., the UCSD-Ped2, CHUCK Avenue, and ShanghaiTech. This demonstrates the model’s superior performance in detecting anomalies compared to state-of-the-art methods, showcasing its potential to significantly enhance automated video surveillance systems by achieving area under the receiver operating characteristic curve (AUC ROC) values of 95.6, 86.8, and 82.1. To show the effectiveness of the proposed framework in detecting anomalies in extra-large datasets, we trained the model on a subset of the huge contemporary CHAD dataset that contains over 1 million frames, achieving AUC ROC values of 71.8 and 64.2 for CHAD-Cam 1 and CHAD-Cam 2, respectively, which outperforms the state-of-the-art techniques.

  • Research Article
  • 10.4314/jobasr.v4i2.29
Optimized deep learning and KNN Models with PCA feature selection for forecasting Cowpea yeild in Nigeria
  • Apr 14, 2026
  • Journal of Basics and Applied Sciences Research
  • Terfa Benjamin Yecho + 1 more

Reliable prediction of crop yield plays a critical role in improving agricultural decision-making and promoting sustainable farming practices. Conventional approaches are often limited in their ability to model the nonlinear relationships that exist among plant growth dynamics. In this study, three machine learning techniques namely; Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM) networks and K-Nearest Neighbors (KNN) were designed and assessed for their effectiveness in forecasting cowpea (Vigna unguiculata) yield based on data acquired from IoT-enabled sensors, the dataset was obtained from a controlled cultivation experiment, with yield quantified by the number of pods produced per plant. Principal Component Analysis reduced dimensionality while preserving over 95% of the variance. Hyperparameter optimization was performed using GridSearchCV for KNN and Keras Tuner RandomSearch for CNN and LSTM. The optimized achieved the highest predictive accuracy with an R² of 0.8754, MAE of 0.0492, and RMSE of 0.0639, outperforming the optimized LSTM (96 units, additional dense layer of 64 units, RMSprop optimizer; R² = 0.8334) and optimized KNN (n_neighbors = 3; R² = 0.7566). However, both CNN and LSTM showed systematic under-prediction bias at higher yield values, with LSTM exhibiting the largest negative residuals. The findings demonstrate that deep learning approaches, particularly CNN, can effectively model crop yield with relatively small datasets when combined with appropriate feature selection and hyperparameter optimization. The integration of statistical feature selection with agronomic domain knowledge enhances model robustness and biological interpretability. These results support improved decision-making in precision agriculture, enabling more accurate yield forecasting for sustainable cowpea production.

  • Conference Article
  • Cite Count Icon 1
  • 10.1117/12.2652301
Anomaly detection and recognition of video surveillance images based on deep learning
  • Nov 10, 2022
  • Zehao Bao + 3 more

In order to solve the problems of pool generalization ability of traditional algorithms and high cost of manual inspection for abnormal image detection in remote video surveillance, this paper proposes an algorithm for abnormal image detection in video surveillance based on deep learning. First, the convolutional neural network based on VGG-16 uses the he_normal method to initialize the weights, and then the self-made datasets is preprocessed and input into the convolutional neural network for training, and finally an image for detecting video surveillance is obtained Model of abnormal interference. Experimental results show that this method can detect abnormal interference such as overexposure of brightness, color distortion, and video freezes in video surveillance, with an accuracy rate of 86%.

  • Research Article
  • Cite Count Icon 4
  • 10.3390/agriculture15030307
A Model for Diagnosing Mild Nutrient Stress in Facility-Grown Tomatoes Throughout the Entire Growth Cycle
  • Jan 30, 2025
  • Agriculture
  • Yunpeng Yuan + 4 more

The effective diagnosis of mild nutrient stress across the complete growth cycle of facility-grown tomatoes is challenging. This study proposes a deep learning framework based on CNN + LSTM, using canopy near-infrared spectroscopy from different growth stages of tomatoes as input, to diagnose mild stress of nitrogen (N), potassium (K), and calcium (Ca) throughout the entire growth cycle of facility-grown tomatoes. The study compares the diagnostic performance of Random Forest (RF), Support Vector Machine (SVM), Partial Least Squares (PLS), Convolutional Neural Networks (CNNs), and CNN + Long Short-Term Memory (LSTM) models for detecting mild nutrient stress in facility-grown tomatoes. Firstly, the preprocessing method of spectral characteristic bands combined with Savitzky-Golay (SG) + Standard Normal Variate (SNV) was determined. Subsequently, all sample data were divided into six groups: N-deficient, K-deficient, Ca-deficient, N-excess, K-excess, and Ca-excess. The aforementioned models were then used for classification prediction. The results show that RF and CNN + LSTM models demonstrated good predictive performance. Specifically, RF achieved accuracy rates of 70.14%, 90.81%, 88.59%, and 85.37% in the classification tasks of Ca-deficient, N-excess, K-excess, and Ca-excess, respectively. The CNN + LSTM model achieved accuracy rates of 93.33%, 63.33%, 99.2%, 83.33%, and 98.52% in the classification tasks of K-deficient, Ca-deficient, N-excess, K-excess, and Ca-excess, respectively. Finally, in the Leave-One-Group-Out Validation (LOGOV) for validating the model’s generalisation performance, RF performed better in the N-deficient, K-deficient, and Ca-deficient tasks, achieving diagnostic accuracy rates of 80.19%, 81.43%, and 77.02%, respectively. The CNN + LSTM model showed a diagnostic accuracy rate of 66.72% in the N-excess classification task. The study concludes that, given complete training data, the CNN + LSTM model can effectively diagnose mild nutrient stress (N, K, and Ca) in facility-grown tomatoes in most scenarios.

  • Conference Article
  • Cite Count Icon 1
  • 10.52591/lxai2023061810
Anomaly Detection in Surveillance Videos Using Spatio-Temporal Context Information
  • Jun 18, 2023
  • Hernan Benitez-Restrepo + 1 more

Several computer vision algorithms have been proposed to detect anomalous activities (robberies, murders, vandalism, among others) in videos. According to the learning approach, they can be classified into probabilistic distribution modeling, sparse coding, and deep learning based methods. The main drawbacks of these approaches are (i) extraction of low-level features that do not capture complex behaviors of instances on the scene, (ii) generation of features from irrelevant regions, (iii) overlooking of relationships among objects, and (iv) omission of long-term dependencies. To solve these issues, we propose a deep learning architecture that leverages the relationships among objects. It achieves this by using an attention mechanism and learning long-term dependencies using a multilayer recurrent neural network (multilayer LSTM). An AUC score of 0.749 on the UCF-Crime dataset confirms that the proposed algorithm competes effectively against several state-of-the-art approaches for anomaly detection in surveillance videos. It also explains the relationship between regions in the video frames and the anomaly detections.

  • Research Article
  • 10.55041/ijsrem16617
Comparative Analysis of Deep Learning Approaches for Twitter Text Classification
  • Oct 21, 2022
  • INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
  • Lukesh Kadu

Abstract—Sentiment analysis (also known as opinion mining or emotion AI) is the use of natural language processing, text analysis, computational linguistics, and biometrics to systematically identify, extract, quantify, and study affective states and subjective information. Sentiment analysis is widely applied to voice of the customer materials such as reviews and survey responses, online and social media, and healthcare materials for applications that range from marketing to customer service to clinical medicine. With the rise of deep language models, such as RoBERTa, also more difficult data domains can be analyzed, e.g., news texts where authors typically express their opinion/sentiment less explicitly. Sentiment analysis aims to extract opinion automatically from data and classify them as positive and negative. Twitter widely used social media tools, been seen as an important source of information for acquiring people’s attitudes, emotions, views, and feedbacks. Within this context, Twitter sentiment analysis techniques were developed to decide whether textual tweets express a positive or negative opinion. In contrast to lower classification performance of traditional algorithms, deep learning models, including Convolution Neural Network (CNN) and Bidirectional Long Short-Term Memory (Bi-LSTM), have achieved a significant result in sentiment analysis. Keras is a Deep Learning (DL) framework that provides an embedding layer to produce the vector representation of words present in the document. The objective of this work is to analyze the performance of deep learning models namely Convolutional Neural Network (CNN), Simple Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM), bidirectional Long Short-Term Memory (Bi-LSTM), BERT and RoBERTa for classifying the twitter reviews. From the experiments conducted, it is found that RoBERTa model performs better than CNN and simple RNN for sentiment classification. Keywords—Convolution Neural Network (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Deep Learning, Bidirectional Long Short-Term Memory (BiLSTM), Bidirectional Encoder Representations from Transformers (BERT), Robustly Optimized BERT Pre-training Approach (RoBERTa).

  • Conference Article
  • Cite Count Icon 38
  • 10.1109/icip.2018.8451816
Foreground Detection in Surveillance Video with Fully Convolutional Semantic Network
  • Oct 1, 2018
  • Chuming Lin + 2 more

Foreground detection is an important part of surveillance video analysis, and also has challenges. For example, the classical methods are difficult to distinguish the foreground, which is similar to the background. In recent years, Convolutional Neural Networks (CNNs) have been widely used in image processing and achieved better performance. In this paper, we proposed an efficient deep Fully Convolutional Semantic Networks (FCSN) model for foreground detection in surveillance video. Our model aimed at learning the global differences between the video frame and the background image, and the semantic information by utilizing the pre-trained weights on semantic segmentation. In the experiment, unlike other related work, we proposed a reasonable method, which is able to avoid overfitting results to construct training data with 20 videos and test data with 6 videos on the dataset of 2014 ChangeDetection.net (CDnet 2014). Experimental results verified that our model outperforms the state-of-the-art methods in the foreground detection of surveillance video.

  • Research Article
  • 10.52783/pmj.v34.i3.1778
Anomaly Detection in Video Surveillance: A Comparative Analysis of Deep Learning Models
  • Oct 1, 2024
  • Panamerican Mathematical Journal
  • Sangita Mahendra Rajput

Anomaly detection in video surveillance is critical for enhancing security and public safety across various applications, including traffic monitoring, public spaces, and industrial settings. Traditional methods often struggle with the complexity and variability of real-world data, prompting a shift towards advanced machine learning models. This paper presents a comprehensive analysis of deep learning algorithms, including YOLOv5, 3D CNNs, LSTM, Deep SVDD, Vision Transformers, Temporal Transformers, and Autoencoders, applied to three benchmark datasets: CIFAR-10, MVTec AD, and UCSD Anomaly Detection. We compare these algorithms based on accuracy, precision, recall, and F1-score, providing insight into their strengths and weaknesses. The results suggest that Vision Transformers and CNN-LSTM hybrids offer superior performance across spatial and temporal anomaly detection tasks.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/icecct.2019.8869360
A Study on the use of State-of-the-Art CNNs with Fine Tuning for Spatial Stream Generation for Activity Recognition
  • Feb 1, 2019
  • Mercy Ranjit + 1 more

Recurrent neural network (RNN) models have been proven successful in modeling the temporal dynamics in videos of which Long Short-Term Memory (LSTM) networks have been specifically successful as it does not suffer from the vanishing gradient problem. They along with Convolutional Neural Networks (CNN) for visual feature extraction are popularly referred as the Long-term Recurrent Convolutional Networks (LRCN) and have been widely accepted in the recent times for activities like video activity classification, video captioning and video description. The features for these models may be generated using single spatial stream or dual streams, both spatial and motion streams from the video frames. The paper is a study on how the State-of-the-Art networks like ResNet50, InceptionV3 and MobileNet perform with fine tuning for spatial feature extraction in the task of activity recognition in videos using LRCN with stacked LSTM. The fine-tuning approach and optimization settings for the extraction of the visual features from the State-of-the-Art pretrained networks is also discussed in this paper.

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface