Multi-Label Emotion Recognition of Korean Speech Data Using Deep Fusion Models

Seoin Park,Seunghyun Lee,Byeonghoon Jeon,Janghyeok Yoon

doi:10.3390/app14177604

Abstract

As speech is the most natural way for humans to express emotions, studies on Speech Emotion Recognition (SER) have been conducted in various ways However, there are some areas for improvement in previous SER studies: (1) while some studies have performed multi-label classification, almost none have specifically utilized Korean speech data; (2) most studies have not utilized multiple features in combination for emotion recognition. Therefore, this study proposes deep fusion models for multi-label emotion classification using Korean speech data and follows four steps: (1) preprocessing speech data labeled with Sadness, Happiness, Neutral, Anger, and Disgust; (2) applying data augmentation to address the data imbalance and extracting speech features, including the Log-mel spectrogram, Mel-Frequency Cepstral Coefficients (MFCCs), and Voice Quality Features; (3) constructing models using deep fusion architectures; and (4) validating the performance of the constructed models. The experimental results demonstrated that the proposed model, which utilizes the Log-mel spectrogram and MFCCs with a fusion of Vision-Transformer and 1D Convolutional Neural Network–Long Short-Term Memory, achieved the highest average binary accuracy of 71.2% for multi-label classification, outperforming other baseline models. Consequently, this study anticipates that the proposed model will find application based on Korean speech, specifically mental healthcare and smart service systems.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Multi-Label Emotion Recognition of Korean Speech Data Using Deep Fusion Models

Abstract

Talk to us

Similar Papers

More From: Applied Sciences

Lead the way for us

Journal: Applied Sciences	Publication Date: Aug 28, 2024
License type: CC BY 4.0

Similar Papers

Fusion of mel and gammatone frequency cepstral coefficients for speech emotion recognition using deep C-RNN
U Kumaran ... Senthil Murugan Nagarajan
International Journal of Speech Technology | VOL. 24
U Kumaran, et. al.U Kumaran ... Senthil Murugan Nagarajan
13 Jan 2021
International Journal of Speech Technology | VOL. 24

An uncertainty-aware self-training framework with consistency regularization for the multilabel classification of common computed tomography signs in lung nodules.
Ketian Zhan ... Lingxiao Zhou
Quantitative imaging in medicine and surgery | VOL. 13
Ketian Zhan, et. al.Ketian Zhan ... Lingxiao Zhou
01 Sep 2023
Quantitative imaging in medicine and surgery | VOL. 13

E‐Speech: Development of a Dataset for Speech Emotion Recognition and Analysis
Wenjin Liu ... Haoming Liu
International Journal of Intelligent Systems | VOL. 2024
Wenjin Liu, et. al.Wenjin Liu ... Haoming Liu
01 Jan 2024
International Journal of Intelligent Systems | VOL. 2024

Two-Way Sign Language Translator for Visually Impaired People
Ms Kavitha R
INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT | VOL. 08
Ms Kavitha RMs Kavitha R
20 Mar 2024
INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT | VOL. 08

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Multi-Label Emotion Recognition of Korean Speech Data Using Deep Fusion Models

Abstract

Talk to us

Similar Papers

More From: Applied Sciences