Multiplex graph aggregation and feature refinement for unsupervised incomplete multimodal emotion recognition

Yuanyue Deng,Jintang Bian,Shisong Wu,Jian-Huang Lai,Xiaohua Xie

doi:10.1016/j.inffus.2024.102711

Abstract

Multimodal Emotion Recognition (MER) involves integrating information of various modalities, including audio, visual, text and physiological signals, to comprehensively grasp human sentiments, which has emerged as a vibrant area within human–computer interaction. Researchers have developed many methods for this task, but many of these methods rely on labeled supervised learning and struggle to address the issue of missing some modalities of data. To address these issues, we propose a Multiplex Graph Aggregation and Feature Refinement framework for unsupervised incomplete MER, comprising four modules: Completion, Aggregation, Refinement, and Embedding. Specifically, we first capture the correlation information between samples using the graph structures, which aids in the completion of missing data and the multiplex aggregation of multimodal data. Then, we perform refinement operations on the aggregated features as well as alignment and enhancement operations on the embedding features to obtain the fused feature representations, which are consistent, highly separable and conducive to emotion recognition. Experimental results on multimodal emotion recognition datasets demonstrate that our method achieves state-of-the-art performance among unsupervised methods, validating its effectiveness.

Full Text