Articles published on Facial expression recognition
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
5253 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.visres.2026.108827
- Jul 1, 2026
- Vision research
- Yi-Fan Li + 3 more
Using neural networks to understand static and dynamic cues in facial expression recognition.
- New
- Research Article
- 10.1038/s41598-026-60089-6
- Jun 29, 2026
- Scientific reports
- Achraf Jallaglag + 3 more
Video-based emotion recognition is an important topic in affective computing, with applications in human-computer interaction, mental health, and multimedia systems. In this work, we propose a two-stage deep learning approach that extracts spatial features from individual frames and leverages temporal consistency across video sequences for facial expression recognition using the RAVDESS dataset. First, videos are split into frames, and a fine-tuned VGG16 CNN extracts discriminative spatial features from each frame. Second, these features are aggregated into sequences and the frame-level predictions are aggregated at the video level using a majority voting strategy to ensure temporal consistency. Our experiments show that the proposed method achieves 93.6% accuracy at the frame level and 98.1% at the video level, outperforming baseline models and remaining competitive with recent state-of-the-art approaches. Temporal aggregation helps reduce misclassifications of subtle emotions, while fine-tuning improves feature extraction. The approach is computationally efficient and provides a solid foundation for future research in multimodal emotion recognition and advanced video-level aggregation.
- New
- Research Article
- 10.1080/02699931.2026.2693720
- Jun 26, 2026
- Cognition and Emotion
- Simone Favelle + 2 more
ABSTRACT Facial expression recognition depends on both motion and intensity, but how these factors interact within and across emotions remains unclear. Here we examined recognition accuracy of basic emotions (anger, disgust, fear, joy, sadness, surprise) across low, medium, and high intensity displays of static and naturally dynamic facial expressions. Results showed while there was an overall increase in accuracy with intensity, there were complex interactions between intensity, motion, and emotion. Dynamic cues enhanced recognition of joy, surprise, and anger, particularly at low intensity, whereas disgust, fear, and sadness were recognised with greater accuracy from static expressions. We found common confusion patterns of fear-surprise, anger-sadness, and disgust-anger. However, contrary to reports of equivalent confusion patterns across dynamic and static expressions, motion reduced confusions of surprise as fear, but increased confusions of fear as surprise. This novel finding suggests that temporal unfolding facilitates recognition of surprise, but obscures cues aiding identification of fear. These patterns show that the benefits of dynamic information are intensity and emotion-specific and limited when emotions share morphological features. The current findings highlight the complexity of the role of motion and intensity in facial expression recognition and underscore the importance of employing naturally unfolding dynamic stimuli.
- New
- Research Article
- 10.1186/s40359-026-04966-9
- Jun 23, 2026
- BMC psychology
- Busra Izgi + 4 more
Emotion recognition (ER) refers to the perceptual and cognitive processes involved in identifying emotional expressions in others, whereas alexithymia reflects difficulties in identifying and describing one's own emotions. Mood and anxiety symptoms, as well as exposure to stress, have been associated with alterations in emotional processing. This study examined the associations between stress levels, self-emotion processing (alexithymia), and recognition of others' facial expressions in a non-clinical student sample. One hundred twenty-four college students completed questionnaires assessing stress, anxiety (BAI), depression (BDI), and alexithymia (TAS-20). Stress typology measurements included Childhood Trauma Questionnaire (CTQ), Perceived Stress Scale (PSS) and Chronic Stress Scale (CSS). The Emotion Recognition (ER-40) task of PennCNB assessed the ability to recognize angry, fearful, sad, happy, and neutral facial expressions. ER-40 performance was not significantly correlated with stress measures or psychopathology scores, except for the negative correlation of recognition of fear expressions with TAS-20 scores. The relatively high accuracy levels and limited variability observed in task performance may have reduced sensitivity to detect subtle associations, suggesting the possibility of ceiling effects. In contrast, all stress and psychopathology measures were positively correlated with alexithymia scores (BAI: p < 0.001, BDI: p < 0.001, CTQ: p = 0.008, PSS: p < 0.001, CSS: p = 0.001). In linear regression analysis, alexithymia, particularly "difficulty in identifying feelings" subscale (TAS_DIF), was found to be associated with scores in CSS and PSS, when corrected for age, gender and CTQ. Mediation analysis indicated that the association between stress measures and TAS_DIF was statistically associated through anxiety symptoms (p < 0.001), while associations through depressive symptoms were not significant. Higher levels of perceived and chronic stress were associated with greater difficulties in identifying one's own emotions, and these associations were statistically mediated in part by anxiety symptoms. No significant associations were observed between stress and facial emotion recognition performance in this sample, nevertheless, the null findings related to the ER-40 should be interpreted cautiously, particularly in light of potential ceiling effects and the relatively high-functioning nature of the sample. Interventions targeting anxiety and stress management may be beneficial and should be studied for students presenting with elevated stress and alexithymic traits, potentially supporting emotional clarity and adaptive coping in preventive mental health settings.
- New
- Research Article
- 10.1038/s41467-026-74769-4
- Jun 23, 2026
- Nature communications
- Yongbiao Zhai + 9 more
The dynamic vision sensor captures visual information as discrete events, enabling high-speed imaging with reduced data redundancy, but is limited by lack of color sensitivity, a speed-noise trade-off, and inefficient data transfer. Here we show an amphibian-inspired dynamic vision system (ADVS) based on ferroelectric field-effect transistors that emulates the hierarchical functions of amphibian retinas, including spectral perception, spatial preprocessing, and event-driven neural encoding. The ferroelectric transistors exhibit broadband photosensitivity (365-637 nm) and bidirectional photoresponses, enabling multichannel spectral recognition from ultraviolet to visible light. Device arrays further reproduce center-surround receptive-field-like processing, enhancing spatial contrast while suppressing background noise under weak illumination. Owing to the steep switching characteristics (SSmin = 53.8 mV dec-1) of the transistors, the system also supports microsecond-scale event-driven spiking responses. Combined with bioinspired hierarchical preprocessing framework and event-driven convolutional neural network, the ADVS achieves 96.5% accuracy in dynamic facial expression recognition and real-time multi-agent trajectory prediction with <5% error.
- New
- Research Article
- 10.48175/ijarsct-37161
- Jun 22, 2026
- International Journal of Advanced Research in Science Communication and Technology
- Darekar Pranav Pravin And Prof Dr Satish N Gujar
It is essential to study the emotions experienced by the audience in response to advertisements in order to analyze the effectiveness of advertising and engagement of the audience. The conventional way to measure these emotions relies on the use of surveys and feedback forms which may be cumbersome and may not always measure the emotions experienced by the audience. This research work seeks to employ deep learning based approach for the identification of emotions of the students based on their facial expressions while watching advertisements. The system will rely on the usage of CNNs and transfer learning techniques in order to detect emotions from the facial expressions. These emotions are further combined with other engagement parameters in order to compute the multimodal engagement scores indicating the level of interest and engagement of the students. The objective of the system is to classify facial expressions including happiness, sadness, anger, surprise, fear and neutral emotions using the deep learning models trained on the facial patterns from images and video frames. Abstract In today's digital age, the viewer's emotional response plays a vital role in measuring the effectiveness of advertisements. The existing methods to analyse viewer's response include questionnaires, interviews, and manual feedback; however, the methods cannot always determine real-time emotions of viewers. Emotional responses are usually reflected in one's facial expressions, which make Facial Expression Recognition (FER) an excellent method to automatically analyse viewer's emotions.
- New
- Research Article
- 10.1038/s41598-026-53853-1
- Jun 21, 2026
- Scientific reports
- Mauricio Priego + 5 more
Facial Expression Recognition (FER) has become increasingly relevant in engineering applications spanning security, healthcare, human-robot interaction, and intelligent systems. Accurate recognition of emotional states through facial cues is particularly important for enabling autonomous systems to interpret human intent in dynamic environments. This study presents a survey of recent advances in FER, with a focus on machine learning (ML) and deep learning (DL) methodologies developed between 2019 and 2024. Special attention is given to Facial Action Unit Detection (FAUD) as an intermediate and physiologically grounded representation for expression recognition. The analyzed literature is organized according to the final scope of the proposed models, distinguishing FAUD-oriented, FER-oriented, AU-assisted FER, and FAUD-to-FER approaches. Through this structured comparison, the survey identifies methodological trends, reported performance patterns, and the conditions under which FAUD-based intermediate modeling appears particularly advantageous for reliable FER.
- Research Article
- 10.1016/j.jbtep.2026.102120
- Jun 17, 2026
- Journal of behavior therapy and experimental psychiatry
- Marta Bodecka-Zych + 5 more
Shaping new perceptions: A preliminary multi-method investigation of changes in hostile attributions following a psychoeducational mentalization-based treatment module.
- Research Article
- 10.1088/1361-6501/ae739f
- Jun 5, 2026
- Measurement Science and Technology
- Cao Shiming
A hybrid feature attention network with transformer integration for facial expression recognition in the wild
- Research Article
- 10.1186/s13023-026-04427-x
- Jun 3, 2026
- Orphanet journal of rare diseases
- Pei-Zhu Zhang + 9 more
Emotional states are known to modulate tremor severity in Wilson's disease (WD), but the neural mechanisms underlying this emotion-tremor coupling remain poorly understood. This study aimed to investigate the neurophysiological and structural substrates of emotion-induced tremor variability in WD patients using electroencephalography (EEG) microstate analysis, kinematic tremor tracking, and structural MRI. Forty-five tremor-dominant WD patients and 20 healthy controls underwent assessments with the Body Image Disturbance Questionnaire (BIDQ), Fahn-Tolosa-Marin Tremor Rating Scale (FTM-TRS), Self-Assessment Manikin (SAM), and Facial Expression Recognition (FER) tasks. Tremor kinematics were quantified via Kinovea, and EEG microstates were analyzed during emotion induction. Structural magnetic resonance imaging (MRI) was used to evaluate brain atrophy patterns. WD patients exhibited greater social avoidance (Z = -5.721, p < 0.05), prolonged FER reaction times (1.78s vs. 1.05s, p < 0.001), and more negative face selections (6 vs. 5, p = 0.036). Negative emotional states were associated with significantly larger tremor amplitude (3.79 px vs. 3.16 px, p = 0.012). EEG microstates showed increased frequency of microstate C (4.74/min vs. 3.41/min, p < 0.001) and coverage (16.62% vs. 13.22%, p < 0.001) during negative emotion, which correlated with tremor severity (ρ = 0.319, p = 0.039). Regression analysis identified lentiform nucleus damage (β = 0.361), cerebellar atrophy (β = 0.300), and frontal atrophy (β = -0.386) as predictors of emotional arousal (R² = 0.277, p = 0.042). Emotion-tremor coupling in WD involves dysregulated salience networks and cerebellar-frontal-lentiform circuits, with EEG microstate C as a potential marker for targeted interventions.
- Research Article
- 10.1093/cercor/bhag095
- Jun 2, 2026
- Cerebral cortex (New York, N.Y. : 1991)
- Mingkui Yang + 6 more
Emotion perception mostly occurs in social contexts where people engage in emotional communication. Research has revealed that contextual faces affect facial expression perception through 2 distinct ways: congruence and functional relation. However, their robustness under varying levels of cognitive load remains unclear. Thus, this study utilized event-related potentials to explore how cognitive load modulated these effects in facial expression processing. Participants performed a facial expression recognition task under different cognitive load conditions, identifying the emotional expression of a centrally presented target face, while a peripherally presented contextual face displayed a dynamic emotional expression and gazed toward the target. Behavioral results showed fearful faces were recognized faster under low than high cognitive load, but no significant difference for angry faces. Event-related potential results found cognitive load modulated both early and late stages of facial processing. Importantly, when recognizing fearful target faces, late positive potential amplitudes under low cognitive load were more positive after angry than fearful contextual faces (functional-relational context effects). In contrast, under high cognitive load, late positive potential amplitudes were more positive after fearful than angry contextual faces (congruence effects). These findings suggested cognitive load modulated both effects on facial expression processing, with this modulation occurring during late-stage processing.
- Research Article
- 10.1016/j.metip.2026.100241
- Jun 1, 2026
- Methods in Psychology
- L Stacchi + 4 more
Dynamic facial expressions of emotion are typically recognized more accurately than static expressions, especially among individuals with immature, vulnerable or impaired facial expression recognition (FER) systems, such as children, older adults, and individuals with clinical conditions. These findings underscore the need of assessing both dynamic and static stimulus formats and suggest that the dynamic advantage could serve as a potential individual-level marker of FER impairment. However, previous research has primarily focused on group-level effects, often overlooking critical individual differences. Here, we tested whether the QUEST threshold-seeking algorithm can efficiently estimate the minimal signal required for accurate recognition of static and dynamic facial expressions at the individual level. We also quantified the minimum number of trials needed to obtain stable threshold estimates. To evaluate its sensitivity, we compared FER thresholds across neurotypical young adults, older adults, and a well-documented case of acquired prosopagnosia. Our findings demonstrate that the QUEST algorithm is a robust and efficient tool for rapidly estimating meaningful FER thresholds at the individual level. We also provide guidance on the minimum number of trials required to obtain stable threshold estimates and identify the facial expressions that serve as the most sensitive markers for probing the dynamic advantage. This psychophysical approach is particularly well-suited for single-subject analyses, assessment of FER in populations with limited capacities, and inclusion in comprehensive testing batteries. Collectively, this psychophysical approach enables scalable, time-efficient screening and longitudinal monitoring of FER, facilitating cross-individual and population comparisons in both research and clinical settings.
- Research Article
- 10.1002/aur.70249
- Jun 1, 2026
- Autism research : official journal of the International Society for Autism Research
- Mohamad Hassan Fadi Hijab + 4 more
Play holds a prominent role in Autism Research and practice, serving not only as a medium for assessment, providing insights into developmental progress and diagnostic indicators, but also as a foundational element in targeted interventions designed to support specific skills, education, and therapeutic outcomes. Additionally, free, unstructured play offers autistic children vital opportunities for self-expression, exploration, and emotional regulation. Recent advancements in computer vision techniques have significantly expanded the possibilities in Autism Research by capturing nuanced motor behaviors and facilitating interactive environments tailored to enhance social engagement and responsiveness. This paper investigates the application of computer vision techniques within diverse play contexts involving autistic children. A systematic literature review was conducted across six major databases, identifying 39 studies published between January 2014 and January 2026 that met the inclusion criteria. The collected data included the type of computer vision techniques used, participant demographics, and purpose of play and its environment. Most of the included studies used structured play scenarios, employing methods such as facial expression recognition, eye tracking, and pose estimation to objectively assess social communication, emotion recognition, and motor coordination in autistic children. Some studies positioned their approach as targeted interventions, while others utilized play-based tasks for early screening and diagnostic insights. The findings highlight the usefulness of computer vision in enhancing child-centered play and autism-related assessments. However, small sample sizes, methodological inconsistencies, and limited real-world applicability remain significant constraints. Future research should focus on inclusive designs, larger and more diverse participant samples, and longitudinal studies to validate the long-term effectiveness of these systems. By addressing these limitations, computer vision-enhanced play could become a crucial component of both personalized interventions and diagnostic frameworks for autistic individuals.
- Research Article
- 10.1016/j.eswa.2026.131475
- Jun 1, 2026
- Expert Systems with Applications
- Afifa Khelifa + 2 more
Unified framework for in-the-wild recognition of basic and compound facial expressions using label distribution learning and dynamic ensemble selection
- Research Article
- 10.1016/j.gmod.2026.101327
- Jun 1, 2026
- Graphical Models
- Xiaoyu Fan + 6 more
Deep spatial and channel sliding attention patches for pose-invariant facial expression recognition
- Research Article
- 10.1192/bjp.2026.10664
- May 28, 2026
- The British journal of psychiatry : the journal of mental science
- Amy L Gillespie + 10 more
Selective serotonin reuptake inhibitors (SSRIs) are limited by inadequate response in a significant proportion of patients, slow onset, minimal cognitive benefit and side-effects. Preclinical studies suggest selective serotonin 4 receptor (5-HT4R) agonists may produce faster antidepressant effects via distinct mechanisms; however, there has been no experimental research in clinical populations to date. To test whether the novel 5-HT4R partial agonist PF-04995274 produces early behavioural and neural changes in emotional cognition similar to SSRIs in patients with unmedicated major depressive disorder (MDD). In a double-blind, placebo-controlled trial, 90 participants with MDD were randomised to 7 days of PF-04995274 (15 mg), citalopram (20 mg) or placebo. Emotional processing was assessed using a behavioural facial expression recognition task and functional magnetic resonance imaging (fMRI) of implicit emotional face processing (days 6-9). Observer- and self-reported symptoms of depression were also measured at baseline and study end. As anticipated, citalopram reduced relative accuracy and increased relative reaction time to identify negative faces, with corresponding changes in neural activity (reduced left amygdala activation to emotional faces and valence-specific shifts in cortical regions). In contrast, PF-04995274 produced no change in behavioural negative bias or amygdala activity but increased medial-frontal cortex activation across valences. While this was not a clinical trial, both active treatments demonstrated an early treatment response with reduced observer-rated depression severity relative to placebo; PF-04995274 also reduced self-reported depression, state anxiety and negative affect. PF-04995274 did not show the typical antidepressant profile of negative bias reductions observed with citalopram. Instead, it was associated with distinct increased medial-frontal activation during an emotional faces task, coupled with preliminary evidence of early clinical improvement, suggesting a potential alternative pathway for antidepressant effects. Findings support further clinical trials of 5-HT4R agonists and investigation of pro-cognitive and mood effects. NCT03516604.
- Research Article
- 10.1080/01973533.2026.2666504
- May 23, 2026
- Basic and Applied Social Psychology
- Hanyu Zhang + 3 more
Children’s ability to recognize facial expressions plays a critical role in their emotional understanding and cognitive development. Previous research suggests that mindfulness training can effectively enhance children’s facial expression recognition, yet the underlying mechanisms remain unclear, and evidence from real-world educational settings is limited. This study first conducted a cross-sectional analysis to examine the relationship between mindfulness and emotion recognition in preschoolers, with attention level tested as a mediator. A randomized controlled trial was then implemented: the intervention group received a 4-month mindfulness-based curriculum, while the control group engaged in regular non-mindfulness classroom activities. Children’s facial expression recognition and attentional stability were assessed before the intervention, immediately after, and three months post-training using standardized facial recognition and cancelation tasks. Results showed that mindfulness training significantly improved both emotion recognition and attentional stability. Moreover, increases in mindfulness predicted long-term gains in facial expression recognition, with changes in attentional stability fully mediating this relationship. These findings suggest that integrating mindfulness into early childhood education may be helpful for enhancing attentional stability and emotional perception, offering valuable insights for promoting young children’s emotional and social development.
- Research Article
- 10.65521/ijeecs.v15i1s.3074
- May 22, 2026
- International Journal of Electrical, Electronics and Computer Systems
- Samiksha Butle + 4 more
Interview preparation has become increasingly complex due to the growing demand for technical competency, communication proficiency, behavioral intelligence, and real-time decision-making skills. Traditional mock interview methods often depend on human evaluators, resulting in subjective assessment, limited scalability, high operational costs, and inconsistent feedback mechanisms. This survey presents a comprehensive review of Artificial Intelligence (AI)-driven mock interview systems designed to automate and enhance interview preparation through intelligent candidate evaluation. The study explores recent advancements in Natural Language Processing (NLP), speech and emotion analysis, computer vision–based behavioral assessment, and Large Language Model (LLM)-based conversational agents for realistic interview simulation. Based on the analysis of existing approaches, a unified multimodal framework is proposed that integrates textual response evaluation, speech confidence analysis, facial expression recognition, sentiment understanding, and adaptive interview generation. The proposed architecture incorporates explainable AI for transparent feedback, visual analytics for progress monitoring, multilingual capabilities for broader accessibility, and cloud-based deployment for scalability. Additional features such as bias mitigation, learning platform integration, gamification, and secure data management further improve usability and effectiveness. The proposed intelligent framework aims to deliver objective, personalized, scalable, and cost-effective interview readiness assessment while addressing limitations of existing isolated evaluation systems.
- Research Article
- 10.3390/bs16050792
- May 16, 2026
- Behavioral Sciences
- Alessandro De Santis + 4 more
Background. Facial expression recognition depends on how visual information is sampled across the face over time. Static area-of-interest (AOI) measures describe where observers look but provide limited information about the sequential organization of gaze. This study examined how gaze is organized during facial expression recognition and whether this organization remains comparable across two conditions differing in the temporal order of contextual and facial stimuli. Methods. Eye-tracking data were collected from 27 participants performing a facial expression recognition task. Fixations on faces were mapped onto three AOIs: Upper Facial Zone (UFZ), Central Facial Zone (CFZ), and Lower Facial Zone (LFZ). Gaze organization was examined using first- and second-order Markov models, entropy estimates, spatial repositioning measures, and a gaze stability index. Results. Gaze transitions showed a structured, non-random organization centered on the CFZ. In the first-order Markov model, transitions from both the UFZ and LFZ were directed primarily toward the CFZ, and within-zone transitions were also most likely in the CFZ. Entropy was lower for the CFZ than for the upper and lower regions, indicating lower transition uncertainty in the central region. The second-order model showed an influence of recent fixation history while preserving the predominance of the CFZ. Spatial repositioning varied across facial zones in both conditions. However, mixed-effects analyses showed no effect of condition on gaze stability. Conclusions. Facial expression recognition was associated with a pattern of exploration in which the central facial region emerged as the most likely fixation destination, with limited evidence of condition-related differences in gaze organization.
- Research Article
- 10.1371/journal.pone.0348906
- May 15, 2026
- PLOS One
- Ying Lu + 3 more
Under the influence of unsafe emotions, miners’ ability to perceive risks is hindered, which can easily lead to decision-making errors and safety accidents. To recognize unsafe emotions exhibited by miners during operations, this study proposes a deep learning-based bimodal framework that integrates speech and facial expression features. A convolutional neural network (CNN) combined with a bidirectional long short-term memory (Bi-LSTM) network is employed to model local spectral patterns and temporal dependencies in speech signals, and ShuffleNet-V2 is used to capture deep facial features. In addition, three feature enhancement strategies are proposed to improve the generalization ability of the model. By constructing a dataset containing five categories of miners’ unsafe emotions for network training, the model achieves a mean recognition accuracy of 85.56%. Furthermore, we conducted a preliminary field test of the bimodal model in a real mining environment. The results provide preliminary evidence of its potential applicability in real-world mining conditions.