International Workshop on Multimodal Learning - 2023 Theme: Multimodal Learning with Foundation Models
This workshop highlights the transformative impact of foundation models like BERT and GPT-3 on multimodal learning, emphasizing interdisciplinary collaboration to address challenges in data fusion across modalities such as language, vision, and sensors, and fostering new research directions and applications in this rapidly evolving field.
The recent advancements in machine learning and artificial intelligence (particularly foundation models such as BERT, GPT-3, T5, ResNet, etc.) have demonstrated remarkable capabilities and driven significant revolutionary changes to the way we make inferences from complex data. These models represent a fundamental shift in the way data are approached and offer exciting new research directions and opportunities for multimodal learning and data fusion. Given the potential of foundation models to transform the field of multimodal learning, there is a need to bring together experts and researchers to discuss the latest developments in this area, exchange ideas, and identify key research questions and challenges that need to be addressed. By hosting this workshop, we aim to create a forum for researchers to share their insights and expertise on multimodal data fusion and learning using foundation models, and to explore potential new research directions and applications in the rapidly evolving field. We expect contributions from interdisciplinary researchers to study and model interactions between (but not limited to) modalities of language, graphs, time-series, vision, tabular data, sensors, and more. Our workshop will emphasize interdisciplinary work and aim at seeding cross-team collaborations around new tasks, datasets, and models.
- Research Article
53
- 10.1016/j.ins.2022.12.014
- Dec 9, 2022
- Information Sciences
Analysis of multimodal data fusion from an information theory perspective
- Supplementary Content
124
- 10.3390/s20236856
- Nov 30, 2020
- Sensors (Basel, Switzerland)
Multimodal learning analytics (MMLA), which has become increasingly popular, can help provide an accurate understanding of learning processes. However, it is still unclear how multimodal data is integrated into MMLA. By following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, this paper systematically surveys 346 articles on MMLA published during the past three years. For this purpose, we first present a conceptual model for reviewing these articles from three dimensions: data types, learning indicators, and data fusion. Based on this model, we then answer the following questions: 1. What types of data and learning indicators are used in MMLA, together with their relationships; and 2. What are the classifications of the data fusion methods in MMLA. Finally, we point out the key stages in data fusion and the future research direction in MMLA. Our main findings from this review are (a) The data in MMLA are classified into digital data, physical data, physiological data, psychometric data, and environment data; (b) The learning indicators are behavior, cognition, emotion, collaboration, and engagement; (c) The relationships between multimodal data and learning indicators are one-to-one, one-to-any, and many-to-one. The complex relationships between multimodal data and learning indicators are the key for data fusion; (d) The main data fusion methods in MMLA are many-to-one, many-to-many and multiple validations among multimodal data; and (e) Multimodal data fusion can be characterized by the multimodality of data, multi-dimension of indicators, and diversity of methods.
- Research Article
22
- 10.3390/agronomy13020401
- Jan 30, 2023
- Agronomy
Fruit quality is an important aspect in determining the consumer preference in the supply chain. Thermal imaging was used to determine different pineapple varieties according to the physicochemical changes of the fruit by means of the deep learning method. Deep learning has gained attention in fruit classification and recognition in unimodal processing. This paper proposes a multimodal data fusion framework for the determination of pineapple quality using deep learning methods based on the feature extraction acquired from thermal imaging. Feature extraction was selected from the thermal images that provided a correlation with the quality attributes of the fruit in developing the deep learning models. Three different types of deep learning architectures, including ResNet, VGG16, and InceptionV3, were built to develop the multimodal data fusion framework for the classification of pineapple varieties based on the concatenation of multiple features extracted by the robust networks. The multimodal data fusion coupled with powerful convolutional neural network architectures can remarkably distinguish different pineapple varieties. The proposed multimodal data fusion framework provides a reliable determination of fruit quality that can improve the recognition accuracy and the model performance up to 0.9687. The effectiveness of multimodal deep learning data fusion and thermal imaging has huge potential in monitoring the real-time determination of physicochemical changes of fruit.
- Research Article
742
- 10.1162/neco_a_01273
- May 1, 2020
- Neural Computation
With the wide deployments of heterogeneous networks, huge amounts of data with characteristics of high volume, high variety, high velocity, and high veracity are generated. These data, referred to multimodal big data, contain abundant intermodality and cross-modality information and pose vast challenges on traditional data fusion methods. In this review, we present some pioneering deep learning models to fuse these multimodal big data. With the increasing exploration of the multimodal big data, there are still some challenges to be addressed. Thus, this review presents a survey on deep learning for multimodal data fusion to provide readers, regardless of their original community, with the fundamentals of multimodal deep learning fusion method and to motivate new multimodal data fusion techniques of deep learning. Specifically, representative architectures that are widely used are summarized as fundamental to the understanding of multimodal deep learning. Then the current pioneering multimodal data fusion deep learning models are summarized. Finally, some challenges and future topics of multimodal data fusion deep learning models are described.
- Preprint Article
3
- 10.36227/techrxiv.174137752.20965470/v1
- Mar 7, 2025
Today, remote sensing (RS) can offer earth observation data with various temporal-spatial-spectral characteristics by leveraging different kinds of sensors, forming a multimodal data framework. Shifting the perspective from orbit to ground, text, points of interest (POIs), street-view images and other geospatial data can provide information related to but distinct from RS images. In order to enrich the information capacity and build a comprehensive understanding, multimodal learning in the RS field attempts to simultaneously process and apply data of various modalities. However, multimodal approaches in supervised manner often require expensive human annotation, impeding the full release of data potential. To alleviate this issue, self-supervised learning (SSL) has become an attractive way to learn from unlabeled data, which can extract meaningful representations by designing effective pretext learning objectives. The strengths of label-free, featureextraction and task-agnostic allow SSL to easily scale up the data and model size, paving the way for RS multimodal foundation model (FM). In this survey, we systematically review the evolving field of RS multimodal SSL. In terms of data modalities, this review not only covers research utilizing multimodal RS images but also includes studies that integrate RS images with other forms of geospatial data, providing a comprehensive overview of the data integration scenarios. At the methodology level, multimodal SSL requires the synergy of learning objective and data fusion module. To provide a systematical framework for understanding the trends and challenges of RS multimodal SSL approaches, we present a structured methodology taxonomy in terms of multimodal SSL objective and data fusion strategy. And for each type of method, we summarize its characteristics, key elements and common scenarios. As for the application areas, we categorize them into four classes, including image processing, image understanding, vision-language understanding, and socioeconomic prediction. In addition, we also provide a systematical review of RS multimodal FMs based on SSL. Finally, we discuss challenges and future directions of RS multimodal SSL. It is our aspiration that this review will act as a starting point for researches to examine the advancements and engage in the exploration of RS multimodal SSL studies.
- Book Chapter
33
- 10.1007/978-3-030-32251-9_69
- Jan 1, 2019
Effective fusion of multi-modality neuroimaging data, such as structural magnetic resonance imaging (MRI) and fluorodeoxyglucose positron emission tomography (PET), has attracted increasing interest in computer-aided brain disease diagnosis, by providing complementary structural and functional information of the brain to improve diagnostic performance. Although considerable progress has been made, there remain several significant challenges in traditional methods for fusing multi-modality data. First, the fusion of multi-modality data is usually independent of the training of diagnostic models, leading to sub-optimal performance. Second, it is challenging to effectively exploit the complementary information among multiple modalities based on low-level imaging features (e.g., image intensity or tissue volume). To this end, in this paper, we propose a novel Deep Latent Multi-modality Dementia Diagnosis (DLMD\(^2\)) framework based on a deep non-negative matrix factorization (NMF) model. Specifically, we integrate the feature fusion/learning process into the classifier construction step for eliminating the gap between neuroimaging features and disease labels. To exploit the correlations among multi-modality data, we learn latent representations for multi-modality data by sharing the common high-level representations in the last layer of each modality in the deep NMF model. Extensive experimental results on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset validate that our proposed method outperforms several state-of-the-art methods.
- Research Article
- 10.1038/s41598-026-36296-6
- Jan 20, 2026
- Scientific reports
Multimodal Machine Learning (MML) methods address various efficient ways of driving insights from various data modalities, e.g., in healthcare settings, tabular electronic health records along with other modalities, such as medical imaging, electrocardiogram data (ECG), and textual doctors' notes and reports. Using deep learning methods, we propose a novel MML approach for mortality prediction in healthcare settings that fuses tabular data, ECG, and written notes in various stages. To this end, this research addresses various challenges related to MML including (1) collecting and building comprehensive data representations from various modalities that may require different preprocessing steps to handle noise and distorted data, (2) ensuring data alignment across modalities, and (3) choosing the optimal fusion strategy (i.e., early, late, or hybrid). This study uses three distinct data modalities: tabular data (encompassing healthcare records, vital signs in real-time, laboratory test results, procedures, and diagnosis records), ECG data, and textual notes from doctors about patients. These modalities are obtained from the MIMIC-IV, MIMIC-ECG, and MIMIC-IV-Note datasets, which include comprehensive medical records, ECG reports, and textual doctors' notes to explore and evaluate methods in all MML stages. The methodology includes data preprocessing to address noise, outliers, and missing values. It involves comparing fusion strategies (early, late, hybrid) for integrating multimodal data. In addition, novel deep learning models that use attention mechanisms are implemented for better data interaction. Model performance is evaluated with metrics like AUC-ROC, precision, recall, and F-score. The results of our proposed multimodal neural network model using multimodal information showed a substantial increase in performance, with an AUC of 0.96, surpassing the performance of previous single modality literature models. Using multimodal data, the aim is to make the proposed model obtain a holistic view of patient health similar to that of domain experts, resulting in better informed clinical decisions and potentially better clinical outcomes. Our promising results suggest the need to examine biases in training data, such as mortality class imbalances, to improve model performance. Future work should also address the interpretability of complex deep learning models for clinical adoption.
- Dissertation
- 10.32657/10356/182226
- Jan 1, 2024
In the past few years, multimodal learning has made significant progress. The goal of multimodal learning is to create models that can relate and process data from various modalities. One of the challenges is to learn useful representations efficiently given the heterogeneity of the data. Another is how to fuse the information from two or more modalities to perform a prediction, which is robust against possibly missing modalities. To reduce these research gaps, this dissertation attempts to develop effective and efficient network modules for both unimodal learning and crossmodal fusion. It also aims to improve the robustness of the fused features for different downstream tasks. In multimodal representation learning, both complementary crossmodal representation fusion and effective unimodal representation are crucial. Some prior works try to modulate one modal feature to another directly. Although it can be effective in aligning the multimodal features, it will ignore both unimodal and crossmodal representation refinements, which is important for multimodal fusion. In this dissertation, we introduce the Unimodal and Crossmodal Refinement Network (UCRN) to enhance both unimodal and crossmodal representations in multimodal learning. We propose a unimodal refinement module that iteratively updates modality-specific representations using transformer-based attention layers, followed by self-quality improvement layers. These refined unimodal representations are then projected into a common latent space and further tuned using a crossmodal refinement module. The results in multiple benchmark datasets show improved performance and robustness against missing modalities and noisy data in multimodal sequence fusion scenarios. Besides representation refinement for better fusion performance, it is also important to reduce the overfitting issue during learning. As the predictive powers between modalities are different, the existing modality gap can lead to overfitting and undermine the fusion performance. This dissertation aims to improve unimodal and crossmodal representations by the proposed regularized expressive representation distillation (RERD) approach. To improve crossmodal optimization and minimize modality gaps before fusion, a multimodal Sinkhorn distance regularizer is introduced, and multi-head distillation encoders with iterative updates are used to refine unimodal representations. We evaluate the proposed method on a range of benchmark datasets. The results show that RERD performs better than current baselines, proving to be an effective method for deep multimodal fusion on sequence datasets. To further improve the robustness of multimodal representations against noisy inputs, we study the robustness in the context of multimodal contrastive learning (MCL), as contrastive learning is effective at discriminating coexisting semantic features (positive) from irrelative ones (negative) in multimodal signals. To address weakness in MCL, this dissertation presents Pace-adaptive and Noise-resistant Noise-Contrastive Estimation (PN-NCE) as a novel self-supervised method for multimodal fusion. We propose to adaptively optimize the similarity between positive and negative pairs and improve robustness against noisy inputs during training. By integrating an estimator to measure modality invariance, PN-NCE achieves consistent performance improvements across various multimodal tasks and datasets and comparable results with supervised learning approaches. To gain more insight into effective and reliable multimodal learning in practical applications, we examine the proposed method of audio-visual deception detection in videos. Deception detection in conversations is a challenging yet important task, having pivotal applications in various fields. The first challenge is the scarcity of high-quality datasets in deception detection research. In this dissertation, we introduce a large gameshow deception detection dataset, DOLOS, with rich multimodal annotations. DOLOS comprises 1,675 video clips with audio-visual annotations featuring 213 subjects. We benchmark deception detection approaches on the DOLOS dataset. Additionally, we propose Parameter-Efficient Crossmodal Learning (PECL), where we propose a Uniform Temporal Adapter and a Plug-in Audio-Visual Fusion module, to enhance performance with fewer parameters and exploit multi-task learning for improved deception detection performance. The Uniform Temporal Adapter module is different from the previous ones in UCRN and RERD because it is lightweight and plug-and-play. In summary, this dissertation focuses on efficient and robust multimodal learning and fusion. To achieve these goals, different methods and modules are proposed to enhance the performance of fused features for downstream tasks. Experimental results on different benchmark datasets and real-world applications show the effectiveness of the proposed method compared with state-of-the-art approaches.
- Research Article
2
- 10.7507/1001-5515.202310011
- Oct 25, 2024
- Sheng wu yi xue gong cheng xue za zhi = Journal of biomedical engineering = Shengwu yixue gongchengxue zazhi
Currently, the development of deep learning-based multimodal learning is advancing rapidly, and is widely used in the field of artificial intelligence-generated content, such as image-text conversion and image-text generation. Electronic health records are digital information such as numbers, charts, and texts generated by medical staff using information systems in the process of medical activities. The multimodal fusion method of electronic health records based on deep learning can assist medical staff in the medical field to comprehensively analyze a large number of medical multimodal data generated in the process of diagnosis and treatment, thereby achieving accurate diagnosis and timely intervention for patients. In this article, we firstly introduce the methods and development trends of deep learning-based multimodal data fusion. Secondly, we summarize and compare the fusion of structured electronic medical records with other medical data such as images and texts, focusing on the clinical application types, sample sizes, and the fusion methods involved in the research. Through the analysis and summary of the literature, the deep learning methods for fusion of different medical modal data are as follows: first, selecting the appropriate pre-trained model according to the data modality for feature representation and post-fusion, and secondly, fusing based on the attention mechanism. Lastly, the difficulties encountered in multimodal medical data fusion and its developmental directions, including modeling methods, evaluation and application of models, are discussed. Through this review article, we expect to provide reference information for the establishment of models that can comprehensively utilize various modal medical data.
- Research Article
43
- 10.1016/j.dibe.2023.100198
- Jul 13, 2023
- Developments in the Built Environment
Multimodal integration for data-driven classification of mental fatigue during construction equipment operations: Incorporating electroencephalography, electrodermal activity, and video signals
- Research Article
433
- 10.1038/s41746-022-00712-8
- Nov 7, 2022
- NPJ digital medicine
Machine learning is frequently being leveraged to tackle problems in the health sector including utilization for clinical decision-support. Its use has historically been focused on single modal data. Attempts to improve prediction and mimic the multimodal nature of clinical expert decision-making has been met in the biomedical field of machine learning by fusing disparate data. This review was conducted to summarize the current studies in this field and identify topics ripe for future research. We conducted this review in accordance with the PRISMA extension for Scoping Reviews to characterize multi-modal data fusion in health. Search strings were established and used in databases: PubMed, Google Scholar, and IEEEXplore from 2011 to 2021. A final set of 128 articles were included in the analysis. The most common health areas utilizing multi-modal methods were neurology and oncology. Early fusion was the most common data merging strategy. Notably, there was an improvement in predictive performance when using data fusion. Lacking from the papers were clear clinical deployment strategies, FDA-approval, and analysis of how using multimodal approaches from diverse sub-populations may improve biases and healthcare disparities. These findings provide a summary on multimodal data fusion as applied to health diagnosis/prognosis problems. Few papers compared the outputs of a multimodal approach with a unimodal prediction. However, those that did achieved an average increase of 6.4% in predictive accuracy. Multi-modal machine learning, while more robust in its estimations over unimodal methods, has drawbacks in its scalability and the time-consuming nature of information concatenation.
- Conference Article
20
- 10.1109/igarss47720.2021.9554255
- Jul 11, 2021
With the ever-growing availability of different remote sensing (RS) products from both satellite and airborne platforms, simultaneous processing and interpretation of multimodal RS data have shown increasing significance in the RS field. Different resolutions, contexts, and sensors of multimodal RS data enable the identification and recognition of the materials lying on the earth's surface at a more accurate level by describing the same object from different points of the view. As a result, the topic on multimodal RS data fusion has gradually emerged as a hotspot research direction in recent years. This paper aims at presenting an overview of multimodal RS data fusion in several mainstream applications, which can be roughly categorized by 1) image pansharpening, 2) hyperspectral and multispectral image fusion, 3) multimodal feature learning, and (4) crossmodal feature learning. For each topic, we will briefly describe what is the to-be-addressed research problem related to multimodal RS data fusion and give the representative and state-of-the-art models from shallow to deep perspectives.
- Research Article
- 10.1109/msp.2025.3604570
- Sep 1, 2025
- IEEE signal processing magazine
Multimodal fusion provides significant benefits over single modality analysis by leveraging both shared and complementary information across diverse data sources. In this article, we systematically review methods for fusion of heterogonous multimodal biomedical data of varying dimensionality (including neuroimaging, biomics, clinical phenotypes and text), with a focus on neuroscience. We discuss the strengths and limitations of these strategies based on a survey of 302 research articles. Next, we examine the applications of these methods to a variety of scenarios spanning a continuum from scientific research to clinical practice. Finally, an in-depth discussion of common challenges and promising directions for future development of multimodal biomedical data fusion are provided. Overall, multimodal fusion shows substantial benefits and transformative potential in the field of neuroscience. Future research should prioritize improving model generalization, enhancing interpretability, addressing inherent data limitations, and developing unified platforms alongside multimodal foundational models to bridge the gap between fusion techniques, research, and application to various domains.
- Conference Article
1
- 10.1109/ictbig59752.2023.10456100
- Dec 8, 2023
Parkinson's disease is a neurological ailment that affects millions of individuals throughout the world. Successful treatment will greatly enhance the quality of life for those who have it. To effectively treat Parkinson's disease, this study investigates how multimodal data and machine learning technology might be combined. To improve illness diagnosis, symptom severity prediction, and medication efficacy evaluation, the suggested technique draws from a broad variety of data sources such as clinical records, medical imaging, and sensor data. In this study, novel data fusion and analytic techniques are used to generate a comprehensive picture of the patient's health. Classification and prediction can be accomplished using either the Random Forest, Convolutional Neural Network, or Long Short-Term Memory Network machine learning techniques. Data fusion and feature extraction, the first steps in building reliable models, are graphically shown by a set of mathematical equations. Data presented in the tables of comparison show how much of an improvement the proposed remedy represents over the status quo. This strategy has the potential to significantly alter the way doctors treat Parkinson's disease, leading to better outcomes for patients and setting a new standard for healthcare in general.
- Research Article
64
- 10.1109/embc.2019.8856500
- Jul 1, 2019
- Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference
Early prediction of diseased brain conditions is critical for curing illness and preventing irreversible neuronal dysfunction and loss. Generically regarding the different neuroimaging modalities as filtered, complementary insights of brain's anatomical and functional organization, multimodal data fusion could be hypothesized to enhance the predictive power as compared to a unimodal prediction of disease progression. More recently, deep learning (DL) based methods on structural MRI (sMRI) data have outperformed classical machine learning approaches in several neuroimaging applications including diagnostic classification and prediction. Similarly, functional MRI (fMRI) features estimated using a dynamic (i.e. time-varying) functional connectivity (FC) approach have been found to be more discriminative and predictive of the clinical diagnosis than those based on the static FC approach. Motivated by this, we introduce a novel multimodal data fusion framework featuring deep residual learning of non-linear sMRI features and dynamic FC (dFC) based extraction of fMRI features to predict the subset of individuals with mild cognitive impairments who would progress to Alzheimer's disease within a time-period of three years from the baseline scanning sessions. Our cross-validated results from the developed multimodal (sMRI-fMRI) data fusion framework demonstrate a significant improvement in performance over the unimodal prediction analyses with the fMRI (p = 7.03 x 10-7) and sMRI (p = 6.72 x 10-4) modalities. As such, the findings in this work highlight the benefits of combining multiple neuroimaging data modalities via data fusion, corroborate the predictive value of the tested DL and dFC features and argue in favor of exploration of similar approaches to learn neuroanatomical and functional alterations in the neuroimaging data.