- New
- Research Article
- 10.1016/j.asoc.2026.115212
- Jul 1, 2026
- Applied Soft Computing
- Kavya Johny + 4 more
- New
- Research Article
- 10.1016/j.asoc.2026.115167
- Jul 1, 2026
- Applied Soft Computing
- Mengqi Gao + 3 more
- New
- Research Article
- 10.1016/j.asoc.2026.115289
- Jul 1, 2026
- Applied Soft Computing
- Jaeuk Jang + 1 more
- New
- Research Article
- 10.1016/j.asoc.2026.115202
- Jul 1, 2026
- Applied Soft Computing
- Ertugrul Ayyildiz + 2 more
- New
- Research Article
- 10.1016/j.asoc.2026.115270
- Jul 1, 2026
- Applied Soft Computing
- Hazim Albedran + 4 more
- New
- Research Article
- 10.1016/j.asoc.2026.115112
- Jul 1, 2026
- Applied Soft Computing
- Qilong Xue + 4 more
- New
- Research Article
- 10.1016/j.asoc.2026.115153
- Jul 1, 2026
- Applied Soft Computing
- Antonio Manuel Gómez-Orellana + 6 more
- New
- Research Article
- 10.1016/j.asoc.2026.114998
- Jul 1, 2026
- Applied Soft Computing
- Mohammad Hassan Tayarani Najaran + 3 more
Automatic emotion recognition plays a critical role in areas such as mental-health monitoring, human–robot interaction, and personalised learning systems, yet current multimodal approaches often struggle with high intra-class variability and the limited discriminative power of raw audio–visual features. Existing methods typically rely on direct classification of audio or facial data, which does not explicitly enforce a structured joint embedding in which emotional categories become separable. This paper addresses this limitation by proposing a supervised contrastive feature-mapping algorithm that transforms temporal audio and video features into a representation that minimises intra-class distances while maximising inter-class distances. In contrast to prior work, which usually focuses on handcrafted feature engineering or end-to-end classifiers, our approach explicitly learns a discriminative metric space that enhances the geometry of the feature distribution. The method is evaluated on the RAVDESS and CREMA-D benchmark datasets. Experimental results show that the proposed mapping yields consistent accuracy improvements over strong machine-learning baselines, with gains of up to approximately 6%, achieving 96.07% accuracy on RAVDESS and competitive performance on CREMA-D, while outperforming or matching recent state-of-the-art multimodal emotion-recognition pipelines. Statistical tests (Kruskal–Wallis and paired t-tests) confirm that the learned representation significantly increases class separability ( ). While the method assumes the availability of paired audio–visual inputs without requiring explicit temporal alignment, the learned feature space is compact, discriminative, and well suited to downstream tasks such as affect-aware dialogue systems, rehabilitation monitoring, and adaptive educational interfaces. These results demonstrate that contrastive feature mapping provides a robust and generalisable framework for multimodal emotion analysis. • We propose a supervised contrastive model that maps temporal audio-visual features into a joint representation, minimizing intra-class variations to enhance the separability of emotional classes in the transformed feature space. • We provide an in-depth analysis of feature distributions before and after transformation, demonstrating improved class discrimination. • Our algorithm achieves superior performance compared to existing approaches, and we publicly release the software for further research and development.
- New
- Research Article
- 10.1016/j.asoc.2026.115252
- Jul 1, 2026
- Applied Soft Computing
- Francesco Nitti + 10 more
- New
- Research Article
- 10.1016/j.asoc.2026.115263
- Jul 1, 2026
- Applied Soft Computing
- Haibin Sun + 1 more