EffortNet: A Deep Learning Framework for Objective Assessment of Speech Enhancement Technologies Using EEG-Based Alpha Oscillations.
This paper presents EffortNet, a novel deep learning framework for decoding listening effort at the individual level from electroencephalography (EEG) during speech comprehension. Quantifying listening effort remains a significant challenge in speech-hearing research. We collected 64-channel EEG data from 122 participants during speech comprehension under four conditions: clean, noisy, MMSE-enhanced, and Transformer-enhanced speech. Statistical analyses confirmed that alpha power (8-13 Hz) was significantly higher during noisy speech processing compared with clean or enhanced conditions, confirming its validity as an objective biomarker of listening effort. To address the substantial inter-individual variability in EEG signals, EffortNet integrates three complementary learning paradigms: self-supervised learning to leverage unlabeled data, incremental learning for progressive adaptation to individual characteristics, and transfer learning for efficient knowledge transfer to new subjects. Our experimental results demonstrate that EffortNet achieves ${80}.{9}$ % classification accuracy with only 40% of the training data from new subjects, significantly outperforming conventional CNN ( ${62}.{3}$ %) and STAnet ( ${61}.{1}$ %) models. The probability-based metric derived from our model revealed that Transformer-enhanced speech elicited neural responses more similar to clean speech than MMSE-enhanced speech. This neural pattern aligns with objective metrics, but contrasts with subjective intelligibility ratings. This dissociation underscores the value of objective neural markers in evaluating hearing technologies, such as headsets, hearables, and assistive listening devices.
- Research Article
134
- 10.1007/s10916-019-1486-z
- Dec 13, 2019
- Journal of Medical Systems
Depression or Major Depressive Disorder (MDD) is a mental illness which negatively affects how a person thinks, acts or feels. MDD has become a major disease affecting millions of people presently. The diagnosis of depression is questionnaire based and is not based on any objective criteria. In this paper, feature extracted from EEG signal are used for the diagnosis of depression. Alpha, alpha1, alpha2, beta, delta and theta power and theta asymmetry was used as feature. Alpha1, alpha2 along with theta asymmetry was also used as a feature. Multi-Cluster Feature Selection (MCFS) was used for feature selection when feature combination was used. The classifiers used were Support Vector Machine (SVM), Logistic Regression (LR), Naïve-Bayesian (NB) and Decision Tree (DT). Alpha2 showed higher classification accuracy than alpha1 and alpha power in all applied classifier. From t-test it was found that there was a significant difference in the theta power of left and right hemisphere of normal subjects, but there was no significant difference in depression patients. Average theta asymmetry in normal subjects is higher than MDD patients but the difference in theta asymmetry in normal subjects and MDD patients is not significant. The combination of alpha2 and theta asymmetry showed the highest classification accuracy of 88.33% in SVM.
- Research Article
11
- 10.3390/app14178091
- Sep 9, 2024
- Applied Sciences
In this research, five systems were developed to classify four distinct motor functions—forward hand movement (FW), grasp (GP), release (RL), and reverse hand movement (RV)—from EEG signals, using the WAY-EEG-GAL dataset where participants performed a sequence of hand movements. During preprocessing, band-pass filtering was applied to remove artifacts and focus on the mu and beta frequency bands. The initial system, a preliminary study model, explored the overall framework of EEG signal processing and classification, utilizing time-domain features such as variance and frequency-domain features such as alpha and beta power, with a KNN model for classification. Insights from this study informed the development of a baseline system, which innovatively combined the common spatial patterns (CSP) method with continuous wavelet transform (CWT) for feature extraction and employed a GoogLeNet classifier with transfer learning. This system classified six unique pairs of events derived from the four motor functions, achieving remarkable accuracy, with the highest being 99.73% for the GP–RV pair and the lowest 80.87% for the FW–GP pair in intersubject classification. Building on this success, three additional systems were developed for four-way classification. The final model, ML-CSP-OVR, demonstrated the highest intersubject classification accuracy of 78.08% using all combined data and 76.39% for leave-one-out intersubject classification. This proposed model, featuring a novel combination of CSP-OVR, CWT, and GoogLeNet, represents a significant advancement in the field, showcasing strong potential as a general system for motor imagery (MI) tasks that is not dependent on the subject. This work highlights the prominence of the research contribution by demonstrating the effectiveness and robustness of the proposed approach in achieving high classification accuracy across different motor functions and subjects.
- Research Article
38
- 10.1109/jiot.2023.3320269
- Mar 1, 2024
- IEEE Internet of Things Journal
Emotions are complex, and people vary greatly in their accuracy in recognizing their own emotions and those of others. With advances in computer science and neuroscience, there is a desire to use automated techniques to help people identify emotions. Bio-electrical signals have been proven effective for emotion detection, but the acquisition of conventional electrocardiogram (ECG) and EEG requires medical-specific equipment, which is very expensive, uncomfortable, and inconvenient due to the large number of electrodes and the hair-covered scalp. In this article, a novel emotion recognition method based on the feature fusion of single-lead EEG and ECG signals is proposed, using the long short term memory (LSTM)-MLP-based model and the CNN-based model for feature fusion and classification, respectively, with fivefold cross-validation for validation. The ECG and EEG signals of 15 participants were collected in five states: 1) happy; 2) relaxed; 3) calm; 4) sad; and 5) afraid, each of which was stimulated using the participants’ own proposed music. Various time-domain features, frequency-domain features, and nonlinear features were extracted from the ECG and EEG signals. Experimental results demonstrate that the accuracy of emotion recognition and classification of signals captured by the proposed device can reach 92.08% using the CNN model. While using the LSTM-MLP feature fusion model, the accuracy figure can be improved to 95.07%. The results of the ablation experiment indicate that the feature fusion approach does improve the accuracy of recognition. It is demonstrated that the proposed device and emotional recognition approach are effective and feasible.
- Research Article
48
- 10.1016/j.cortex.2022.02.017
- Mar 19, 2022
- Cortex
Better speech-in-noise comprehension is associated with enhanced neural speech tracking in older adults with hearing impairment
- Research Article
27
- 10.3390/s22020680
- Jan 16, 2022
- Sensors (Basel, Switzerland)
Motion classification can be performed using biometric signals recorded by electroencephalography (EEG) or electromyography (EMG) with noninvasive surface electrodes for the control of prosthetic arms. However, current single-modal EEG and EMG based motion classification techniques are limited owing to the complexity and noise of EEG signals, and the electrode placement bias, and low-resolution of EMG signals. We herein propose a novel system of two-dimensional (2D) input image feature multimodal fusion based on an EEG/EMG-signal transfer learning (TL) paradigm for detection of hand movements in transforearm amputees. A feature extraction method in the frequency domain of the EEG and EMG signals was adopted to establish a 2D image. The input images were used for training on a model based on the convolutional neural network algorithm and TL, which requires 2D images as input data. For the purpose of data acquisition, five transforearm amputees and nine healthy controls were recruited. Compared with the conventional single-modal EEG signal trained models, the proposed multimodal fusion method significantly improved classification accuracy in both the control and patient groups. When the two signals were combined and used in the pretrained model for EEG TL, the classification accuracy increased by 4.18–4.35% in the control group, and by 2.51–3.00% in the patient group.
- Research Article
- 10.1016/j.neuroimage.2026.121866
- May 1, 2026
- NeuroImage
Electroencephalography signals in a female Fragile X Syndrome mouse model.
- Research Article
6
- 10.1016/j.patrec.2023.01.011
- Jan 19, 2023
- Pattern Recognition Letters
Differentiable Mean Opinion Score Regularization for Perceptual Speech Enhancement
- Research Article
3
- 10.1080/25742442.2024.2395217
- Jul 2, 2024
- Auditory Perception & Cognition
Listening effort (LE) is critical to understanding speech perception in acoustically challenging environments. Electroencephalograpy (EEG) alpha power has emerged as a potential neural correlate of LE. However, the magnitude and direction of the relationship between acoustic challenge and alpha power has been inconsistent in the literature. In the current study, a secondary analysis of previously collected dates, we examine the broadband 1/f-like exponent and offset of the EEG power spectrum as measures of aperiodic neural activity during effortful speech perception and the influence of this aperiodic activity on reliable estimation of periodic (i.e. alpha) neural activity. EEG was continuously recorded during sentence listening and the broadband (1–40 Hz) EEG power spectrum was computed for each participant for quiet and noise trials separately. Using the specparam algorithm, we decomposed the power spectrum into both aperiodic and periodic components and found that broadband aperiodic activity was sensitive to background noise during speech perception and additionally impacted the measurement of noise-induced changes on alpha oscillations. We discuss the implications of these results for the LE and neural speech processing literatures.
- Research Article
- 10.1097/aud.0000000000001785
- Feb 1, 2026
- Ear and hearing
Listening fatigue seems to reduce the quality of life for people with hearing loss. However, it has mostly been measured with subjective questionnaires. The present study investigates neural oscillatory activity in the resting electroencephalogram as a marker of the development of listening fatigue. It was hypothesized that alpha and theta power and subjective fatigue would increase after, compared with before, a challenging listening task, and that alpha peak frequency would decrease post-task. Before and after completing an approximately hour-long task of listening to speech-in-noise, resting electroencephalogram in eyes-closed and eyes-open conditions and subjective fatigue ratings were collected from older adults with age-appropriate hearing abilities and healthy cognitive status (n = 41; 29 female; mean age: 69.22 years; range 50 to 85). Paired t tests were used to analyze pre- to post-task changes, and correlations were used to explore relations among variables. Subjective fatigue on a visual-analogue scale increased post-task, indicating that the listening task induced the experience of fatigue. Alpha power increased post-task, whether eyes were open or closed. By contrast, numerical trends for theta power to increase post-task were not significant. In addition, peak alpha frequency decreased post-task. Exploratory analyses indicated a relation of the post-task increases in subjective fatigue and alpha power with eyes closed, and a relation of age and hearing ability to the extent of post-task reduction in peak alpha frequency. Disaggregating results by sex indicated stronger, more consistent trends for women than for men, both for the main effects of time-on-task on alpha power and for relations between oscillatory power and subjective fatigue. Alpha power and peak alpha frequency appear to be promising objective markers of the development of mental fatigue induced by challenging listening tasks in older adults.
- Research Article
7
- 10.1016/j.apacoust.2021.108539
- Dec 4, 2021
- Applied Acoustics
Enhancing the correlation between the quality and intelligibility objective metrics with the subjective scores by shallow feed forward neural network for time–frequency masking speech separation algorithms
- Research Article
3
- 10.1111/bjet.13535
- Nov 18, 2024
- British Journal of Educational Technology
Cognitive load is a critical internal state associated with learners' learning process and significantly influences learning outcomes. With the worldwide popularity of video‐based learning (VBL), tracking real‐time cognitive load variations becomes more and more important for the timely provision of adaptive learning support during the learning process. This study proposed and validated an electroencephalogram (EEG)‐supported approach to tracking real‐time cognitive load variations during continuous VBL. We recruited 108 healthy adult participants to watch a specially designed video lecture with a sequence of interconnected slides of equal length. EEG signals were continuously recorded throughout the session. The video lecture was designed with varying levels of content difficulty (ie, rated from 1 to 5) across slides and was narrated at three different speeds (ie, slow, normal and fast) to induce cognitive load variations. For each slide, the cognitive load was quantified using both subjective ratings (ie, self‐reported difficulty) and an EEG‐derived measure (ie, alpha power). Through linear mixed model analysis, we demonstrated the feasibility of using alpha power to track real‐time cognitive load variations during the continuous VBL process after controlling the effect of mental fatigue. This study provides a foundation for developing learning enhancement technologies that enable the timely provision of adaptive learning support in VBL. Practitioner notesWhat is already known about this topic Video‐based learning has become a prevailing learning method for the current generation. Tracking the internal learning state of learners is essential for the timely provision of adaptive learning support during the video‐based learning process. Cognitive load is a critical aspect of internal learning state. While EEG has been proven to be valuable in assessing average cognitive load of a task, few studies have investigated the feasibility of utilizing EEG to track real‐time cognitive load variations in a task. What this paper adds An EEG‐supported approach was proposed to track real‐time cognitive load variations in video‐based learning. A high consistency was found between subjective ratings and EEG‐derived measure of cognitive load. The presence of mental fatigue exerted a significant impact on EEG‐derived measure of cognitive load. Implications for practice and/or policy Generative AI can be leveraged to facilitate mass production of lectures required in the approach. Real‐time tracking of cognitive load variations in video‐based learning enables the timely provision of adaptive learning supports. Additional research is warranted to mitigate the effect of mental fatigue on real‐time tracking of cognitive load variations.
- Research Article
2
- 10.1121/1.2021096
- Nov 1, 1983
- The Journal of the Acoustical Society of America
We describe a model of speech perception in which lawful variability in the speech signal is treated as a source of additional information, rather than as noise. Excitatory and inhibitory interactions among nodes for phonetic features, phonemes, and words are used to account for aspects of coarticulation, as well as the interaction of bottom-up and top-down processes in perception of speech. Two results from a working computer simulation of this model are presented. First, we show how the simulation is able to process real (digitized) speech and retune the detection of features and phonemes in accord with the context. Second, we demonstrate how perceptual behavior which appears to be guided by rules can be induced without any explicit rules in the system. [Work supported by the Office of Naval Research.]
- Research Article
40
- 10.3758/s13414-018-1635-3
- Nov 30, 2018
- Attention, Perception, & Psychophysics
Talker and listener sex in speech processing has been largely unknown and under-appreciated to this point, with many studies overlooking the possible influences. In the current study, the effects of both talker and listener sex on speech intelligibility were assessed. Different methodological approaches to measuring intelligibility (percent words correct vs. subjective rating scales) and collecting data (laboratory vs. crowdsourcing) were also evaluated. Findings revealed that, regardless of methodology, the spoken productions of female talkers were overall more intelligible than the spoken productions of male talkers; however, substantial variability across talkers was observed. Findings also revealed that when data were collected in the lab, there was an interaction between talker and listener sex. This interaction between listener and talker sex was not observed when subjective ratings were crowdsourced from listener subjects across the USA via Amazon Mechanical Turk, although overall ratings remained similar. This possibly suggests that subjective intelligibility ratings may be vulnerable to bias, and such biases may be reduced by recruiting a more heterogeneous subject pool. Many studies in speech perception do not account for these talker, listener, and methodology effects. However, the present results suggest that researchers should carefully consider these effects when assessing speech intelligibility in different conditions, and when comparing findings across studies that have used different subject demographics and/or methodologies.
- Research Article
2
- 10.1088/2057-1976/aab29a
- Jul 1, 2018
- Biomedical Physics & Engineering Express
Objective. Auditory brain-computer interfaces (BCIs) have gained attention recently due to their potential applicability to severely disabled individuals without functional vision. However, auditory BCIs currently achieve lower accuracies than their visual counterparts, and most have exploited only one type of brain measurement. Recent evidence suggests that the combination of electrical and optical measurements can enhance classification accuracies in motor imagery-based BCIs. The potential of this bimodal combination for auditory BCIs remains unexplored. Approach. We investigated the complementarity of near-infrared spectroscopy (NIRS) and electroencephalography (EEG) in discriminating between attentive and non-attentive behaviours to auditory stimuli. Simultaneous NIRS and EEG signals were recorded from 11 typically developed participants while performing an auditory oddball streaming task. Main Results. Considering neural responses to oddballs during the entire span of a trial, average classification accuracies of 77.43+/−9.6% and 80.7+/−9.5% were achieved using combined EEG and NIRS signal and the EEG-only signal respectively. The combined EEG-NIRS classification accuracy was significantly lower than the EEG-only accuracy in five participants. However, when considering EEG responses to the first oddball within each trial, we found that the inclusion of NIRS activities from the complete task period significantly improved classification accuracies for 2 participants. Significance. Our findings suggest that the consolidation of EEG (midline) and NIRS (temporo-parietal) signals offers limited value beyond an EEG-exclusive approach when decoding prolonged auditory attention-demanding tasks. Future efforts should focus on identifying the optimal time window for the analysis of EEG and NIRS signals to delineate conditions under which the BCI may afford a practical advantage over the EEG-exclusive approach.
- Conference Article
22
- 10.1109/qomex.2012.6263839
- Jul 1, 2012
The typical procedure for evaluating the performance of different objective quality metrics and indices involves comparisons between subjective quality ratings and the quality indices obtained using the objective metrics in question on the known video sequences. Several correlation indicators can be employed to assess how well the subjective ratings can be predicted from the objective values. In this paper, we give an overview of the potential sources for uncertainties and inaccuracies in such studies, related both to the method of comparison, possible inaccuracies in the subjective data, as well as processing of subjective data. We also suggest some general guidelines for researchers to make comparison studies of objective video quality metrics more reliable and useful for the practitioners in the field.