Vision + X: A Survey on Multimodal Learning in the Light of Data.

Ye Zhu,Yu Wu,Yan Yan,Nicu Sebe

doi:10.1109/tpami.2024.3420239

Abstract

We are perceiving and communicating with the world in a multisensory manner, where different information sources are sophisticatedly processed and interpreted by separate parts of the human brain to constitute a complex, yet harmonious and unified sensing system. To endow the machines with true intelligence, multimodal machine learning that incorporates data from various sources has become an increasingly popular research area with emerging technical advances in recent years. In this paper, we present a survey on multimodal machine learning from a novel perspective considering not only the purely technical aspects but also the intrinsic nature of different data modalities. We analyze the commonness and uniqueness of each data format mainly ranging from vision, audio, text, and motions, and then present the methodological advancements categorized by the combination of data modalities, such as Vision+Text, with slightly inclined emphasis on the visual data. We investigate the existing literature on multimodal learning from both the representation learning and downstream application levels, and provide an additional comparison in the light of their technical connections with the data nature, e.g., the semantic consistency between image objects and textual descriptions, and the rhythm correspondence between video dance moves and musical beats. We hope that the exploitation of the alignment as well as the existing gap between the intrinsic nature of data modality and the technical designs, will benefit future research studies to better address a specific challenge related to the concrete multimodal task, prompting a unified multimodal machine learning framework closer to a real human intelligence system.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Vision + X: A Survey on Multimodal Learning in the Light of Data.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on pattern analysis and machine intelligence

Lead the way for us

Journal: IEEE transactions on pattern analysis and machine intelligence	Publication Date: Dec 1, 2024
Citations: 1

Similar Papers

Multi-modal deep learning for automated assembly of periapical radiographs
L Pfänder ... F Schwendicke
Journal of Dentistry | VOL. 135
L Pfänder, et. al.L Pfänder ... F Schwendicke
21 Jun 2023
Journal of Dentistry | VOL. 135

Discriminative multi-modal deep generative models
Fang Du ... Rongrong Fei
Knowledge-Based Systems | VOL. 173
Fang Du, et. al.Fang Du ... Rongrong Fei
02 Mar 2019
Knowledge-Based Systems | VOL. 173

Adapt and explore: Multimodal mixup for representation learning
Ronghao Lin ... Haifeng Hu
Information Fusion | VOL. 105
Ronghao Lin, et. al.Ronghao Lin ... Haifeng Hu
28 Dec 2023
Information Fusion | VOL. 105

Explainable multimodal learning for predictive maintenance of steam generators
Duc An Nguyen ... Kamal Medjaher
PHM Society Asia-Pacific Conference | VOL. 4
Duc An Nguyen, et. al.Duc An Nguyen ... Kamal Medjaher
04 Sep 2023
PHM Society Asia-Pacific Conference | VOL. 4

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Vision + X: A Survey on Multimodal Learning in the Light of Data.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on pattern analysis and machine intelligence