Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Survey on speech emotion recognition: Features, classification schemes, and databases

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Survey on speech emotion recognition: Features, classification schemes, and databases

Similar Papers
  • Research Article
  • Cite Count Icon 26
  • 10.1016/j.apacoust.2020.107519
Investigation of multilingual and mixed-lingual emotion recognition using enhanced cues with data augmentation
  • Jul 22, 2020
  • Applied Acoustics
  • S Lalitha + 3 more

Investigation of multilingual and mixed-lingual emotion recognition using enhanced cues with data augmentation

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 8
  • 10.1007/s10462-024-10760-z
Speech emotion recognition systems and their security aspects
  • May 21, 2024
  • Artificial Intelligence Review
  • Itzik Gurowiec + 1 more

Speech emotion recognition (SER) systems leverage information derived from sound waves produced by humans to identify the concealed emotions in utterances. Since 1996, researchers have placed effort on improving the accuracy of SER systems, their functionalities, and the diversity of emotions that can be identified by the system. Although SER systems have become very popular in a variety of domains in modern life and are highly connected to other systems and types of data, the security of SER systems has not been adequately explored. In this paper, we conduct a comprehensive analysis of potential cyber-attacks aimed at SER systems and the security mechanisms that may prevent such attacks. To do so, we first describe the core principles of SER systems and discuss prior work performed in this area, which was mainly aimed at expanding and improving the existing capabilities of SER systems. Then, we present the SER system ecosystem, describing the dataflow and interactions between each component and entity within SER systems and explore their vulnerabilities, which might be exploited by attackers. Based on the vulnerabilities we identified within the ecosystem, we then review existing cyber-attacks from different domains and discuss their relevance to SER systems. We also introduce potential cyber-attacks targeting SER systems that have not been proposed before. Our analysis showed that only 30% of the attacks can be addressed by existing security mechanisms, leaving SER systems unprotected in the face of the other 70% of potential attacks. Therefore, we also describe various concrete directions that could be explored in order to improve the security of SER systems.

  • Research Article
  • Cite Count Icon 104
  • 10.1016/j.apacoust.2023.109492
Emotional speech Recognition using CNN and Deep learning techniques
  • Jun 28, 2023
  • Applied Acoustics
  • C Hema + 1 more

Emotional speech Recognition using CNN and Deep learning techniques

  • Research Article
  • Cite Count Icon 13
  • 10.1016/j.procs.2022.09.345
Multiple Models Fusion for Multi-label Classification in Speech Emotion Recognition Systems
  • Jan 1, 2022
  • Procedia Computer Science
  • Anwer Slimi + 3 more

Multiple Models Fusion for Multi-label Classification in Speech Emotion Recognition Systems

  • Research Article
  • Cite Count Icon 11
  • 10.1016/j.specom.2023.02.001
Development of a speech emotion recognizer for large-scale child-centered audio recordings from a hospital environment
  • Feb 15, 2023
  • Speech Communication
  • Einari Vaaras + 4 more

In order to study how early emotional experiences shape infant development, one approach is to analyze the emotional content of speech heard by infants, as captured by child-centered daylong recordings, and as analyzed by automatic speech emotion recognition (SER) systems. However, since large-scale daylong audio is initially unannotated and differs from typical speech corpora from controlled environments, there are no existing in-domain SER systems for the task. Based on existing literature, it is also unclear what is the best approach to deploy a SER system for a new domain. Consequently, in this study, we investigated alternative strategies for deploying a SER system for large-scale child-centered audio recordings from a neonatal hospital environment, comparing cross-corpus generalization, active learning (AL), and domain adaptation (DA) methods in the process. We first conducted simulations with existing emotion-labeled speech corpora to find the best strategy for SER system deployment. We then tested how the findings generalize to our new initially unannotated dataset. As a result, we found that the studied AL method provided overall the most consistent results, being less dependent on the specifics of the training corpora or speech features compared to the alternative methods. However, in situations without the possibility to annotate data, unsupervised DA proved to be the best approach. We also observed that deployment of a SER system for real-world daylong child-centered audio recordings achieved a SER performance level comparable to those reported in literature, and that the amount of human effort required for the system deployment was overall relatively modest.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/asyu52992.2021.9598956
RMWSaug: Robust Multi-window Spectrogram Augmentation Approach for Deep Learning based Speech Emotion Recognition
  • Oct 6, 2021
  • Shehu Mohammed Yusuf + 4 more

Data scarcity and speech degradation due to environmental noise are two significant issues in the modelling and deployment speech emotion recognition (SER) systems. Deep learning-based SER systems overfits during modelling because of scarce training samples. Although recent attempts to tackle these issues, simultaneously, using data augmentation have yielded promising results, they are not robust enough to handle speech degradation due to real environmental noise. Thus, there is the need to further improve the classification performance of deployed SER systems. This work proposes an SER system based on a novel robust multi-window spectrogram augmentation (RMWSaug) scheme and, transfer learning to handle these aforementioned issues simultaneously. First, the RMWSaug scheme utilizes the concept of multi-window and multi-noise conditioning of clean speech samples to create additional speech spectrograms required for training. Then, pretrained networks are adapted for speech emotion recognition and finetuned with the generated training datasets to develop a model robust to speech degradation due to noise. Thereby, improving the classification performance in the wild. The Interactive Emotional Dyadic Motion Capture (IEMOCAP) database was selected as benchmark dataset for evaluating the proposed SER system. Experimental results show that the proposed SER system outperformed existing methods when deployed in the wild. The proposed SER system can be deployed to predict the emotions of speakers conversing virtually on online platforms.

  • Conference Article
  • Cite Count Icon 10
  • 10.1109/csci.2015.17
I-Vector Algorithm with Gaussian Mixture Model for Efficient Speech Emotion Recognition
  • Dec 1, 2015
  • Joan Gomes + 1 more

Emotions constitute an essential part of our existence as it exerts great influence on the physical as well as mental health of people. Emotions often play the role of a sensitive catalyst, which fosters lively interaction between human beings. Over the past few decades the focus of researchers on study of the emotional content of speech signals, has progressively increased. Many systems have been proposed to make the Speech Emotion Recognition (SER) process more correct and accurate. The objective of our research is to classify speech emotion implementing a comparatively new method-i-vector model. i-vector model has found much success in the areas of speaker identification, speech recognition and language identification. But it has not been much explored in recognition of emotion. This paper discusses the design of a speech emotion recognition system considering three important aspects. Firstly, i-vector model was implemented in processing extracted features for speech representation. Secondly, an appropriate classification scheme was designed using Gaussian Mixture Model (GMM), Maximum A Posteriori (MAP) adaptation and i-vector algorithm. Finally, the performance of this new system was evaluated using emotional speech database. Speech emotions were identified with this novel system and also with a conventional system and results were compared, which proved that our proposed system can identify speech emotions with less error and more accuracy.

  • Conference Article
  • Cite Count Icon 28
  • 10.1109/icnidc.2018.8525706
Feature Fusion of Speech Emotion Recognition Based on Deep Learning
  • Aug 1, 2018
  • Gang Liu + 2 more

Speech emotion recognition (SER) is a hot topic in academia. One of the key issues in improving the performance of SER systems is the choice of speech emotion features. In order to establish a robust speech emotion recognition system, it is essential to select the features which can be a perfect representation of speech emotion attributes. Researchers has done a lot of work, proposed a variety of emotional features and made great progress. Although each kind of features were proven to be effective, most of methods are based on a single type. In this paper, we proposed a method of feature fusion based on deep learning, combining spectral-based features and pitch-based hyper-prosodic features. The experiments show that this method improves the performance of speech emotion recognition system.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 163
  • 10.3390/s20185212
Deep-Net: A Lightweight CNN-Based Speech Emotion Recognition System Using Deep Frequency Features.
  • Sep 12, 2020
  • Sensors
  • Tursunov Anvarjon + 2 more

Artificial intelligence (AI) and machine learning (ML) are employed to make systems smarter. Today, the speech emotion recognition (SER) system evaluates the emotional state of the speaker by investigating his/her speech signal. Emotion recognition is a challenging task for a machine. In addition, making it smarter so that the emotions are efficiently recognized by AI is equally challenging. The speech signal is quite hard to examine using signal processing methods because it consists of different frequencies and features that vary according to emotions, such as anger, fear, sadness, happiness, boredom, disgust, and surprise. Even though different algorithms are being developed for the SER, the success rates are very low according to the languages, the emotions, and the databases. In this paper, we propose a new lightweight effective SER model that has a low computational complexity and a high recognition accuracy. The suggested method uses the convolutional neural network (CNN) approach to learn the deep frequency features by using a plain rectangular filter with a modified pooling strategy that have more discriminative power for the SER. The proposed CNN model was trained on the extracted frequency features from the speech data and was then tested to predict the emotions. The proposed SER model was evaluated over two benchmarks, which included the interactive emotional dyadic motion capture (IEMOCAP) and the berlin emotional speech database (EMO-DB) speech datasets, and it obtained 77.01% and 92.02% recognition results. The experimental results demonstrated that the proposed CNN-based SER system can achieve a better recognition performance than the state-of-the-art SER systems.

  • Research Article
  • Cite Count Icon 68
  • 10.1016/j.dsp.2020.102763
Speech emotion recognition using cepstral features extracted with novel triangular filter banks based on bark and ERB frequency scales
  • May 11, 2020
  • Digital Signal Processing
  • Sugan Nagarajan + 4 more

Speech emotion recognition using cepstral features extracted with novel triangular filter banks based on bark and ERB frequency scales

  • Research Article
  • Cite Count Icon 103
  • 10.1016/j.specom.2019.09.002
Automatic speech emotion recognition using an optimal combination of features based on EMD-TKEO
  • Sep 19, 2019
  • Speech Communication
  • Leila Kerkeni + 5 more

Automatic speech emotion recognition using an optimal combination of features based on EMD-TKEO

  • Research Article
  • 10.1142/s2196888825300029
Speech Emotion Recognition in Arabic Language: A Review
  • Jul 31, 2025
  • Vietnam Journal of Computer Science
  • Houari Horkous

Nowadays, interpreting human emotions through speech has attracted great attention in human–computer interaction and artificial intelligence. Speech emotion recognition (SER) systems have become a significant field of research. SER is one of the interesting directions in speech processing, is to predict the expressed emotional state. SER systems encounter numerous challenges, such as the availability of appropriate emotional databases, the identification of suitable speech features, and the choice of the appropriate classification method. SER systems are mostly implemented in English, French, German, Indian, and Chinese languages. However, SER for the Arabic language is still in the growing phase. In this work, a literature review on the SER in Arabic has been presented in terms of emotional databases, speech features, and classification algorithms. This review contributes to filling the gap in the works on emotion recognition available in the Arabic language and constitutes a valuable resource for researchers in this field.

  • Research Article
  • Cite Count Icon 105
  • 10.1007/s11042-020-09874-7
Deep learning approaches for speech emotion recognition: state of the art and research challenges
  • Jan 2, 2021
  • Multimedia Tools and Applications
  • Rashid Jahangir + 3 more

Speech emotion recognition (SER) systems identify emotions from the human voice in the areas of smart healthcare, driving a vehicle, call centers, automatic translation systems, and human-machine interaction. In the classical SER process, discriminative acoustic feature extraction is the most important and challenging step because discriminative features influence the classifier performance and decrease the computational time. Nonetheless, current handcrafted acoustic features suffer from limited capability and accuracy in constructing a SER system for real-time implementation. Therefore, to overcome the limitations of handcrafted features, in recent years, variety of deep learning techniques have been proposed and employed for automatic feature extraction in the field of emotion prediction from speech signals. However, to the best of our knowledge, there is no in-depth review study is available that critically appraises and summarizes the existing deep learning techniques with their strengths and weaknesses for SER. Hence, this study aims to present a comprehensive review of deep learning techniques, uniqueness, benefits and their limitations for SER. Moreover, this review study also presents speech processing techniques, performance measures and publicly available emotional speech databases. Furthermore, this review also discusses the significance of the findings of the primary studies. Finally, it also presents open research issues and challenges that need significant research efforts and enhancements in the field of SER systems.

  • Conference Article
  • 10.1109/bharat53139.2022.00042
Refined Feature Vectors for Human Emotion Classifier by combining multiple learning strategies with Recurrent Neural Networks
  • Apr 1, 2022
  • K Swetha + 1 more

The speech emotion recognition (SER) system categorizes human emotions based on contextual features. However, it is seriously affected during the signal transmission in which the quality of real­time speech processing is degraded in the SER system. This paper presents refined feature vectors for human emotion classifiers based on multiple learning strategies combined with recurrent neural networks (Refine­HE­RNN). It extracts spatial emotional vectors by observing speech signals for contextual feature dependency through the multiple learning (ML) approach. It computes signal interpretation, emotional cues, and input correction by using the skip connection (SC) module in the residual block of the ML strategy. The fused layer is simple to concentrate derived features that support automatic learning of classifying different human emotions. For experimental purposes, standard IEMOCAP and MSP­IMPROV datasets are considered for proposed method validation. Results convey that the proposed method has significant improvement (in terms of percentage closer to 80% higher than the existing CNN result) in the feature recognition and is flexible for real­time implementation in the SER system. Moreover, it can extend to automatic sensing of human emotion with the help of a light weighted RNN framework.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 19
  • 10.3390/app13169062
Cross-Corpus Training Strategy for Speech Emotion Recognition Using Self-Supervised Representations
  • Aug 8, 2023
  • Applied Sciences
  • Miguel A Pastor + 4 more

Speech Emotion Recognition (SER) plays a crucial role in applications involving human-machine interaction. However, the scarcity of suitable emotional speech datasets presents a major challenge for accurate SER systems. Deep Neural Network (DNN)-based solutions currently in use require substantial labelled data for successful training. Previous studies have proposed strategies to expand the training set in this framework by leveraging available emotion speech corpora. This paper assesses the impact of a cross-corpus training extension for a SER system using self-supervised (SS) representations, namely HuBERT and WavLM. The feasibility of training systems with just a few minutes of in-domain audio is also analyzed. The experimental results demonstrate that augmenting the training set with EmoDB (German), RAVDESS, and CREMA-D (English) datasets leads to improved SER accuracy on the IEMOCAP dataset. By combining a cross-corpus training extension and SS representations, state-of-the-art performance is achieved. These findings suggest that the cross-corpus strategy effectively addresses the scarcity of labelled data and enhances the performance of SER systems.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant