Articles published on Emotional prosody
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
696 Search results
Sort by Recency
- New
- Research Article
- 10.1037/dev0002182
- Jul 1, 2026
- Developmental psychology
- Ping Tang + 7 more
Prosody and semantic content are important cues in emotion perception, though they are not always congruent, as in irony, humor, and insincerity. Adults typically rely on prosody to interpret emotions when cues conflict, while children gradually shift their reliance from semantics to prosody with age, acquiring adult-like cue-weighting strategies during the school years. However, it was unclear whether children learning tonal languages such as Mandarin Chinese show a similar developmental trajectory, as their intensive experience with linguistically meaningful pitch variations (lexical tones) may heighten sensitivity to emotional prosody and potentially accelerate development. We recruited 104 Mandarin-learning 3- to 6-year olds and 27 adult controls. Stimuli included semantically positive or negative utterances produced with happy or sad prosody, creating prosody-semantics incongruent expressions (e.g., semantically positive utterances with sad prosody, and vice versa). Participants judged the speaker's emotion (happy or sad), and the proportion of prosody-based responses was analyzed. Results showed that 3- and 4-year olds displayed no clear preference for either cue, with prosody-based responses near chance level (50%). Five-year olds began to show prosodic preference (though only for positive utterances with sad prosody), while 6-year olds demonstrated robust prosodic preference for both utterance types, reaching adult-like levels. These findings indicate that Mandarin-learning children develop prosodic preference during preschool years, with emerging preference by age 5 and adult-like strength by age 6. The results suggest that cue-weighting strategies in emotion perception may follow a universal developmental trajectory from semantic-dominance to prosody-dominance, with detailed developmental courses modulated by language-specific experience. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
- Research Article
- 10.1038/s44271-026-00475-y
- Jun 15, 2026
- Communications psychology
- Alessandra Piatti + 3 more
Infants of mothers with severe mental illness are at elevated risk of neurodevelopmental difficulties. This study used functional near-infrared spectroscopy (fNIRS) to compare 8-10-month-old infants exposed to severe maternal mental illness (n = 30) and unexposed controls (n = 30) during two experiments assessing brain activation to voice versus non-voice sounds and to sentences with angry, happy, or neutral prosody. Exposed infants showed atypical right-temporal activation, with stronger activation to non-voice than voice stimuli, unlike controls. Socioeconomic status was positively correlated with activation to non-voice and to sentences, independently of emotional prosody. Further analysis indicated that brain response to non-voice stimuli mediated the relation between socioeconomic status and activation to sentences. These findings reveal distinct neural trajectories in infants exposed tosevere maternal mental illness, detectable within the first year of life, and highlight socioeconomic background as a key factor in early brain development.
- Research Article
- 10.1016/j.bandl.2026.105780
- Jun 5, 2026
- Brain and language
- Carolyn D Gershman + 4 more
The neural basis of emotional prosody processing: development of lateralization.
- Research Article
- 10.1523/jneurosci.1854-25.2026
- Jun 3, 2026
- The Journal of neuroscience : the official journal of the Society for Neuroscience
- Anna Seydell-Greenwald + 4 more
In most healthy adult humans, the right ventro-occipital temporal cortex (vOTC) is more involved in face perception than its left-hemisphere counterpart. Several hypotheses link this right-lateralization for face processing to left-lateralization of language and visual word form processing. In its strongest form, this account assumes an encroachment by word form processing into left vOTC territory that would otherwise be dedicated to face processing, causing face processing that would otherwise be bilateral to rely relatively more on right vOTC. By this logic, if visual word recognition came to rely predominantly on the right hemisphere, face processing should become left-lateralized. Alternatively, it would have to share neural territory with words in right vOTC, which might negatively impact behavioral performance due to crowding. Here, we used functional MRI and the Cambridge Face Memory Test in 15 male and female adolescents and young adults with a history of left-hemisphere perinatal stroke (LHPS) who are right-lateralized for language and word form processing and whose vOTC is anatomically and functionally intact in both hemispheres. Comparison with 14 neurotypical controls revealed no significant group differences in strength or lateralization of face activation and no significant differences in face recognition. These results demonstrate that after LHPS, face and word form processing can both be right-lateralized without significant detriment to either. This is reminiscent of the observation that emotional prosody and sentence processing-typically lateralized to the right and left hemisphere, respectively-are both successfully supported by the right hemisphere in the same participant group.
- Research Article
- 10.1007/s10936-026-10262-9
- Jun 2, 2026
- Journal of psycholinguistic research
- Susan Alizadehfard + 2 more
This study offers a novel contribution to psycholinguistic research by examining how facial emotion recognition predicts the perception of emotional prosody in an unfamiliar language, with a specific focus on Persian speakers and age-related effects. While previous studies have often investigated prosody and facial cues separately, the present work integrates these domains to provide new insights into multimodal social cognition. A total of 210 Persian-speaking adults aged 20-65, divided into three 15-year age groups, completed a computer-based task consisting of visual and auditory components. In the auditory section, 30 emotionally intoned sentences in Greek, representing happiness, sadness, anger, fear, and disgust, were presented. In the visual section, participants viewed 20 morphed facial expressions displaying the same emotions at varying intensities. Regression analyses revealed that emotional prosody recognition was significantly predicted by facial emotion recognition (p < .01), and by age only in the group over 50 years old (p < .01). These findings highlight the novelty of linking cross-linguistic prosody processing with facial emotion recognition and suggest that age-related changes in prosody recognition may reflect cognitive changes in later adulthood. The results underscore the predictive role of facial cues in social cognition and provide implications for understanding age-related differences in social communication.
- Research Article
- 10.1016/j.neubiorev.2026.106638
- Jun 1, 2026
- Neuroscience and biobehavioral reviews
- Jurriaan Witteman
Recent work has suggested that emotional prosody is processed in a ventral and a dorsal processing stream similarly to segmental speech perception. Furthermore, it has been proposed that there is right hemispheric lateralization of emotional prosody perception and that task demands and sex modulate activity within the emotional prosody perception network. In the present study, a novel powerful effect size meta-analysis was performed on the neuroimaging literature of stimulus-driven emotional prosody perception. The results show that a much larger network of areas is activated than previously assumed. Furthermore, all areas assumed to be involved in the ventral and dorsal stream are indeed robustly activated across the literature, except for the dorsal premotor area. Additionally, explicit processing of emotional prosody activated additional areas beyond the core network and seems to engage dorsal stream areas more than implicit processing. No lateralized activation of emotional prosody is observed, suggesting that emotional prosody perception is a highly bilateral process. Last, the proportion of females in primary studies was associated with very subtle enhanced neural processing in neural regions assumed to be involved in both early and late stages of emotional prosody perception. The meta-analytic effect size maps obtained can be used for sample size calculations of future neuroimaging studies of emotional prosody perception.
- Research Article
- 10.1037/pag0000990
- May 4, 2026
- Psychology and aging
- Yi Lin + 2 more
Despite consistent reports on age-related declines in emotion perception, existing studies remain unclear about whether aging exerts a direct influence or an indirect effect via other age-related factors. This study examined the direct and indirect pathways through which age influences multisensory emotional speech perception, specifically focusing on the roles of sensory (auditory) and cognitive functions. In 2022-2024, participants (aged 18-82, N = 182) completed two emotional speech perception tests: a cross-channel test pairing prosody with semantics and a cross-modal test that additionally included visual facial expressions. They were also assessed for hearing sensitivity as well as global and specific cognitive functions, including selective attention, working memory, and musical emotion discrimination. A structural equation modeling approach was applied to examine how these age-related auditory and cognitive factors influenced emotional speech perception across different communication channels. In the cross-channel test, age mediated the performance through cognitive functioning and, to a lesser degree, hearing sensitivity, with both pathways stronger for emotional prosody than for semantics. In the cross-modal test, while the indirect effect of age was enhanced and remained most pronounced in prosody compared to the other two channels, it was primarily mediated via cognitive functioning. In addition, age was found to be a significant direct predictor of perceptual performance, especially in the semantic task. Our findings delineate an integrated, dual-pathway model of age-related changes in multisensory emotional speech perception. This involves test-specific direct age effects, especially for semantic processing, and general indirect effects on prosody perception primarily via cognitive functioning and, to a lesser extent, through hearing sensitivity. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
- Research Article
- 10.1016/j.ajp.2026.104958
- May 1, 2026
- Asian journal of psychiatry
- Anqi Zhou + 9 more
Coupling of emotional prosody recognition and expression and their associations with psychiatric symptoms in schizophrenia.
- Research Article
- 10.1016/j.neuroimage.2026.121980
- May 1, 2026
- NeuroImage
- Leonardo Ceravolo + 10 more
Influence of attention mechanisms on cerebellar and basal ganglia activity during vocal emotion decoding.
- Research Article
- 10.1016/j.heares.2026.109601
- Apr 1, 2026
- Hearing research
- Junsheng Hong + 5 more
The neural processing mode and development of emotional prosody in children with cochlear implants.
- Research Article
- 10.22214/ijraset.2026.77563
- Mar 31, 2026
- International Journal for Research in Applied Science and Engineering Technology
- Anshika Saxena + 1 more
Human emotion recognition has become a significant research focus within artificial intelligence due to its growing importance in human–computer interaction, affective computing, and intelligent decision-support systems. Conventional emotion recognition methods have largely relied on unimodal data sources, such as text, speech, or facial expressions. Although effective in controlled settings, unimodal approaches often provide an incomplete and ambiguous understanding of emotional expression, as human emotions are inherently multimodal. This review paper critically examines a dissertation that proposes a deep learning-based multimodal sentiment analysis framework for human emotion detection by integrating textual, acoustic, and facial expression modalities. The reviewed framework employs a Long Short-Term Memory (LSTM)-based architecture to effectively model temporal and contextual dependencies present in multimodal data. Textual information is encoded using embedded word sequences, audio data captures emotional prosody through acoustic features, and visual inputs represent facial expression patterns. These modality-specific features are fused within a unified deep learning framework to perform binary emotion classification. Experimental evaluation using standard performance metrics, including accuracy, precision, recall, F1- score, confusion matrix analysis, and training–validation curves, demonstrates an overall classification accuracy of 82.22 percent, along with balanced precision and recall values. The review highlights the robustness, methodological soundness, and practical relevance of multimodal sentiment analysis, emphasizing its advantages over unimodal approaches and its contribution to the advancement of affective computing research.
- Research Article
- 10.1080/23273798.2026.2650100
- Mar 27, 2026
- Language, Cognition and Neuroscience
- Yue Lin + 3 more
ABSTRACT In this behavioural and electrophysiological study, we examined the effects of emotional prosody on lexicalisation of newly-learned words in an immersive virtual reality (VR) environment. Chinese speakers learned German words that were accompanied by positive, negative, or neutral emotional prosody. Lexicalisation was subsequently assessed in a semantic priming task while brain activity was simultaneously measured through EEG. The results showed that emotional prosody, regardless of valence, slowed responses to newly learned words relative to those learned with neutral prosody. Critically, words learned under positive prosody conditions elicited larger semantic priming effects in both behavioural performance and the late positive component (LPC). Moreover, the magnitude of this effect was significantly correlated with the participants’ self-reported emotional experience during VR learning. Taken together, our findings suggest that in a VR learning environment, positive prosody facilitates lexicalisation of newly learned words, providing empirical support for the Integrated Cognitive-Affective Theory of Learning with Media.
- Research Article
- 10.1044/2026_jslhr-24-00652
- Mar 27, 2026
- Journal of speech, language, and hearing research : JSLHR
- Nityansh Saluja + 5 more
This study investigates vocal emotion perception in Hindi-speaking children with cochlear implants (CIs) and children with normal hearing (NH) using auditory emotion recognition tasks. Whereas previous research has largely focused on accuracy, this study evaluates both accuracy and reaction time (RT), offering a behavioral measure of processing efficiency in recognizing five vocal emotions: happy, sad, angry, fear, and surprise. Twenty children aged 4-12 years participated in the study: 10 with CIs and 10 with NH, matched for hearing age and gender. Fifty emotionally intoned Hindi sentences were validated and presented auditorily, followed by two emotion-specific animated images. Participants selected the image matching the heard emotion, and their accuracy and RTs were recorded using DMDX software. Statistical analyses included two-way analyses of variance (ANOVAs) for normally distributed emotions (sad, angry, and surprise) and Kruskal-Wallis tests for nonnormally distributed emotions (happy, fear), with Bonferroni-corrected post hoc comparisons. Children with CIs showed significantly lower accuracy and longer RTs compared to children with NH. A two-way ANOVA revealed a significant main effect of group on accuracy, F(1, 54) = 6.20, p = .013, η2 = .06; post hoc analysis indicated significantly lower accuracy for the angry emotion in CI users (p = .006). RT analysis also revealed a robust main effect of group, F(1, 54) = 22.71, p < .001, η2 = .30, indicating slower responses in the CI group (mean RT = 5.57 s ± 0.18) than in the NH group (mean RT = 3.62 s ± 0.18). For normally distributed emotions, between-group differences were significant for sad, angry, and surprise (p < .001). For nonnormally distributed emotions, between-group differences were also significant for happy (p = .006) and fear (p < .001). Thus, all five emotions showed significantly longer RTs in the CI group. Hindi-speaking children with CIs demonstrated reduced speed and accuracy in identifying vocal emotions, indicating challenges in processing emotional prosody. These findings highlight the need for integrating explicit training in emotional prosody within auditory rehabilitation programs for pediatric CI users to improve social communication outcomes.
- Research Article
- 10.1002/icd.70100
- Mar 1, 2026
- Infant and Child Development
- Tyler Birse + 3 more
ABSTRACT We examined English‐speaking preschoolers' and adults' attention to emotional prosody in an unfamiliar language when asked to: (a) match emotional prosody with emotional faces; and (b) use emotional prosody to identify a speaker's intended referent. In Experiment 1, 4‐year‐olds ( N = 36, M = 4.16 years; 18 females) and adults ( N = 38, M = 21.18 years; 26 females) matched happy and sad Polish utterances to a corresponding emotional face, as evidenced through pointing decisions. In Experiment 2, adults ( N = 36, M = 20.17 years; 31 females), but not 4‐year‐olds ( N = 36, M = 4.11 years; 18 females), matched the same emotional utterances to objects whose properties signalled an association with happiness or sadness (e.g., intact vs. broken toy). These findings demonstrate that 4‐year‐olds and adults can recognise emotional prosody in an unfamiliar language, however, only adults are successful at extending this information to other kinds of emotion‐relevant decisions.
- Research Article
- 10.3758/s13423-026-02865-z
- Feb 27, 2026
- Psychonomic bulletin & review
- Aíssa M Baldé + 2 more
Why does musical expertise predict enhanced emotion recognition in speech prosody? Evidence for a causal role of music training is weak, and correlations with musical aptitude could reflect basic auditory abilities rather than musicality per se. Here, we tested whether individual differences in basic auditory discrimination account for the music-prosody association. A total of 164 adults completed forced-choice judgments of emotions in prosody and facial expressions, self-reports of musical experience, objective tests of musical ability (melody and rhythm perception), and adaptive psychoacoustic tasks that estimated discrimination thresholds for pitch, duration, loudness, timbre, and backward masking. Both music training and musical ability correlated with better recognition of prosodic but not facial emotions. The training effect was weak, however, and disappeared after controlling for confounding variables, including general cognitive ability. By contrast, musical ability, specifically melody perception, remained associated with prosodic emotion recognition after accounting for training and other covariates. Crucially, psychoacoustic thresholds correlated with both prosody recognition and melody perception. When considered simultaneously, pitch discrimination alone independently predicted prosodic emotion recognition, but melody perception did not. These findings suggest that music training is an artifactual correlate of prosodic emotion recognition, and that basic pitch sensitivity underlies the link between musical ability and emotional prosody.
- Research Article
- 10.1044/2025_aja-25-00190
- Feb 19, 2026
- American journal of audiology
- Zilong Xie
This study examined the extent to which spectral degradation and speaking style (child-directed vs. adult-directed speech) affect emotion recognition from prosodic cues and how these effects are modulated by concurrent tasks involving nonauditory sensory input. Adults with normal hearing completed an emotion recognition task under three conditions: alone (auditory single-task), concurrently with a low-load visual memory task (four identical images), and with a high-load visual memory task (four different images). Stimuli consisted of semantically neutral sentences spoken in five emotions (angry, happy, neutral, sad, and scared) and two speaking styles (child-directed and adult-directed). All sentences were vocoded to simulate spectral degradation. Emotion recognition was assessed using a single-interval, five-alternative, forced-choice paradigm, in which the participants were asked to indicate which of five emotions was associated with each heard sentence. Emotion recognition was significantly reduced for vocoded stimuli, as indicated by lower sensitivity (d') and prolonged reaction times (RTs). Child-directed speech led to better performance than adult-directed speech, although its facilitative effect was reduced under vocoded conditions. Dual-tasking impaired performance, with lower d' values in both dual-task conditions and slower RTs under high-load dual-task conditions. Crucially, dual-task effects did not significantly vary with spectral degradation or speaking style. Top-down cognitive demands from cross-modal dual-tasking and bottom-up stimulus factors, such as spectral degradation and speaking style, independently influence emotion recognition from prosodic cues. These findings provide insight into how cochlear implant users perceive emotional speech in complex, multimodal environments.
- Research Article
- 10.3389/fcomm.2026.1728758
- Feb 6, 2026
- Frontiers in Communication
- Diana Santos + 1 more
This study investigates the perception of emotional prosody across two major varieties of Portuguese—Brazilian (BP) and European (EP)—examining how differences in their intonational systems and cultural backgrounds shape emotion recognition. Building on theoretical frameworks of vocal emotion, an experiment was designed using acoustically controlled, acted stimuli expressing neutrality, happiness, sadness, and anger. Native listeners from both varieties completed a perception task measuring identification accuracy and reaction times. Results revealed a clear emotion hierarchy: neutrality and sadness were recognized most accurately and rapidly, whereas happiness and anger were frequently confused, indicating higher perceptual ambiguity. While both listener groups showed a native-variety advantage, EP participants demonstrated superior overall accuracy, attributed to greater exposure to BP media, and more consistent cue reliance. These findings highlight the interplay between universal acoustic cues, variety-specific phonological structuring, and cultural exposure in shaping emotional speech perception. This research contributes to prosodic typology in Portuguese, cross-variety speech perception, and the integration of linguistic and affective models, with implications for phonological theory, applied technologies, and communication in general.
- Research Article
- 10.1080/23279095.2026.2624602
- Feb 4, 2026
- Applied Neuropsychology: Adult
- Staša Lalatović + 4 more
Objective The study examined emotion recognition (ER) across visual and auditory modalities in patients with temporal lobe epilepsy (TLE) and frontal lobe epilepsy (FLE), and explored associations with perceived social functioning (SF). Method Fifty patients (30 TLE, 20 FLE) and 50 healthy controls (HC) completed tasks assessing recognition of facial emotion and emotional prosody across seven emotions: neutral, happiness, surprise, anger, disgust, fear, and sadness. Patients also completed the Social Functioning subscale of the Quality of Life in Epilepsy Inventory-31 (QOLIE-31) and self-report questionnaires assessing affective symptoms. Results Both TLE and FLE groups exhibited overall ER deficits across modalities compared to HC, with performance varying by emotion. TLE participants showed difficulties in recognizing fear and disgust across both modalities, whereas FLE participants were impaired in auditory recognition of these emotions and visual recognition of fear. Emotion differentiation impairments were relatively comparable across epilepsy types and modalities. Although groups did not differ in their relative performance across modalities, subsequent correlational analyses revealed a modest association between modalities in patients, but not controls. Within the patient group, the only significant association with perceived SF emerged for recognition of neutral prosodic features in the FLE group. Conclusion Individuals with TLE and FLE experience difficulties recognizing emotions from both facial expressions and vocal cues, especially those with negative valence. Limited associations between ER and perceived SF were observed only in FLE patients. The findings underscore the importance of assessing sociocognitive functioning in PWE.
- Research Article
- 10.1038/s42003-026-09625-8
- Feb 2, 2026
- Communications biology
- Pinyuan Hu + 8 more
Emotional prosody (EP) processing is vital for social communication. Seed-based functional connectivity has been widely used to probe its neural basis, yet most studies rely on part of predefined regions, introducing uncertainty and bias. Furthermore, although gender and task type modulate its activation pattern, their network-level impact remains unclear. Using activation network mapping (a network-level analogue of meta-analysis), we identified a unified EP network and delineated its modulation by gender and task types (explicit or implicit). Results showed broader activation networks in females compared to males, regardless of the task type. Moreover, explicit tasks recruited additional frontal and sensorimotor regions beyond implicit tasks, supporting hierarchical processing. We also identified associations with specific receptors and diseases like autism and Alzheimer's. These findings underscore the importance of considering gender and task type effects on emotional processing research and provide a network-level neural mechanism underlying emotional prosody.
- Research Article
- 10.1111/infa.70087
- Feb 1, 2026
- Infancy : the official journal of the International Society on Infant Studies
- Valentina Silvestri + 5 more
The ability to perceive and respond to vocal emotional cues is critical for early social development, guiding infants' interactions with caregivers. Although newborns are believed to preferentially attend to positive prosody, the perceptual mechanisms underlying this response-as well as the degree to which these responses are language-specific-remain poorly understood. This study examined newborns' sensitivity to emotional prosody above and beyond linguistic prosody marking different sentence types (i.e., interrogatives, imperatives, and declaratives) and its relation to maternal postnatal depressive traits. Forty-three newborns were presented with interrogative, imperative, and declarative utterances spoken each with happy, neutral, and sad prosodies while their non-nutritive sucking (NNS) behavior was measured and maternal depressive traits were assessed. Results revealed inhibited NNS behavior in response to sad and neutral utterances compared to happy ones (ps<0.05), suggesting early differential responsiveness to positive vocal affect, independent of the linguistic prosodic contours of the sentence type. Moreover, reduced responsiveness to non-positive prosodies was positively associated with maternal postnatal depressive traits, suggesting that variation in the early affective environment may shape newborns' auditory-emotional sensitivity. These findings highlight the interplay between biological predispositions and environmental influences in early emotional processing, emphasizing the role of early auditory experiences in shaping socio-emotional development.