Cross-modal Factor Analysis Research Articles

Several information fusion methods are developed for increasing the recognition accuracy in multimodal systems. Canonical correlation analysis (CCA), cross-modal factor analysis (CFA) and their kernel versions are known as successful fusion techniques but they cannot digest the data variability. Probabilistic CCA (PCCA) is suggested as a linear fusion method to capture input variability. A new kernel PCCA (KPCCA) is proposed here to capture both the nonlinear correlations of sources and input variability. The functionality of KPCCA decreases when the number of samples, which determines the size of kernel matrix increases. In the conventional fusion methods the latent variables of different modalities are concatenated; consequently, a large-scale covariance matrix with just limited number of samples must be estimated To overcome this drawback, a sparse KPCCA (SKPCCA) is introduced which scarifies the covariance matrix elements at the cost of decreasing its rank. In the final stage of the gradual evolution of KPCCA, a new feature fusion manner is proposed for SKPCCA (FF-SKPCCA) as a second stage fusion. This proposed method unifies the latent variables of two modalities into a feature vector with an acceptable size. Audio-visual databases like M2VTS (for speech recognition) eNTERFACE and RML (for emotion recognition) are applied to assess FF-SKPCCA compared to state-of-the-art fusion methods. The comparative results indicate the superiority of the proposed method in most cases.

Read full abstract

In this paper, we investigate kernel based methods for multimodal information analysis and fusion. We introduce a novel approach, kernel cross-modal factor analysis, which identifies the optimal transformations that are capable of representing the coupled patterns between two different subsets of features by minimizing the Frobenius norm in the transformed domain. The kernel trick is utilized for modeling the nonlinear relationship between two multidimensional variables. We examine and compare with kernel canonical correlation analysis which finds projection directions that maximize the correlation between two modalities, and kernel matrix fusion which integrates the kernel matrices of respective modalities through algebraic operations. The performance of the introduced method is evaluated on an audiovisual based bimodal emotion recognition problem. We first perform feature extraction from the audio and visual channels respectively. The presented approaches are then utilized to analyze the cross-modal relationship between audio and visual features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. The effectiveness of the proposed solution is demonstrated through extensive experimentation.

Read full abstract

Cross-modal Factor Analysis Research Articles

Articles published on Cross-modal Factor Analysis

Kernel Probabilistic Dependent-Independent Canonical Correlation Analysis

Incomplete Cholesky Decomposition based Kernel Cross Modal Factor Analysis for Audiovisual Continuous Dimensional Emotion Recognition

A Hybrid Latent Space Data Fusion Method for Multimodal Emotion Recognition

Multispectral image change detection with kernel cross-modal factor analysis-based fusion of kernels

FF-SKPCCA: Kernel probabilistic canonical correlation analysis

Adaptive missing texture reconstruction method based on kernel cross-modal factor analysis with a new evaluation criterion

Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Cross-modal Factor Analysis Research Articles

Articles published on Cross-modal Factor Analysis

Kernel Probabilistic Dependent-Independent Canonical Correlation Analysis

Incomplete Cholesky Decomposition based Kernel Cross Modal Factor Analysis for Audiovisual Continuous Dimensional Emotion Recognition

A Hybrid Latent Space Data Fusion Method for Multimodal Emotion Recognition

Multispectral image change detection with kernel cross-modal factor analysis-based fusion of kernels

FF-SKPCCA: Kernel probabilistic canonical correlation analysis

Adaptive missing texture reconstruction method based on kernel cross-modal factor analysis with a new evaluation criterion

Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition