Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A survey of multimodal hybrid deep learning for computer vision: Architectures, applications, trends, and challenges

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

A survey of multimodal hybrid deep learning for computer vision: Architectures, applications, trends, and challenges

Similar Papers
  • Research Article
  • Cite Count Icon 404
  • 10.1007/s00371-021-02166-7
A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets
  • Jun 10, 2021
  • The Visual Computer
  • Khaled Bayoudh + 3 more

The research progress in multimodal learning has grown rapidly over the last decade in several areas, especially in computer vision. The growing potential of multimodal data streams and deep learning algorithms has contributed to the increasing universality of deep multimodal learning. This involves the development of models capable of processing and analyzing the multimodal information uniformly. Unstructured real-world data can inherently take many forms, also known as modalities, often including visual and textual content. Extracting relevant patterns from this kind of data is still a motivating goal for researchers in deep learning. In this paper, we seek to improve the understanding of key concepts and algorithms of deep multimodal learning for the computer vision community by exploring how to generate deep models that consider the integration and combination of heterogeneous visual cues across sensory modalities. In particular, we summarize six perspectives from the current literature on deep multimodal learning, namely: multimodal data representation, multimodal fusion (i.e., both traditional and deep learning-based schemes), multitask learning, multimodal alignment, multimodal transfer learning, and zero-shot learning. We also survey current multimodal applications and present a collection of benchmark datasets for solving problems in various vision domains. Finally, we highlight the limitations and challenges of deep multimodal learning and provide insights and directions for future research.

  • Supplementary Content
  • Cite Count Icon 32
  • 10.1093/genetics/iyae161
A review of multimodal deep learning methods for genomic-enabled predictionin plant breeding
  • Nov 5, 2024
  • Genetics
  • Osval A Montesinos-L\Xf3Pez + 9 more

Deep learning methods have been applied when working to enhance the prediction accuracyof traditional statistical methods in the field of plant breeding. Although deep learningseems to be a promising approach for genomic prediction, it has proven to have somelimitations, since its conventional methods fail to leverage all available information.Multimodal deep learning methods aim to improve the predictive power of their unimodalcounterparts by introducing several modalities (sources) of input information. In thisreview, we introduce some theoretical basic concepts of multimodal deep learning andprovide a list of the most widely used neural network architectures in deep learning, aswell as the available strategies to fuse data from different modalities. We mention someof the available computational resources for the practical implementation of multimodaldeep learning problems. We finally performed a review of applications of multimodal deeplearning to genomic selection in plant breeding and other related fields. We present ameta-picture of the practical performance of multimodal deep learning methods to highlighthow these tools can help address complex problems in the field of plant breeding. Wediscussed some relevant considerations that researchers should keep in mind when applyingmultimodal deep learning methods. Multimodal deep learning holds significant potential forvarious fields, including genomic selection. While multimodal deep learning displaysenhanced prediction capabilities over unimodal deep learning and other machine learningmethods, it demands more computational resources. Multimodal deep learning effectivelycaptures intermodal interactions, especially when integrating data from different sources.To apply multimodal deep learning in genomic selection, suitable architectures and fusionstrategies must be chosen. It is relevant to keep in mind that multimodal deep learning,like unimodal deep learning, is a powerful tool but should be carefully applied. Given itspredictive edge over traditional methods, multimodal deep learning is valuable inaddressing challenges in plant breeding and food security amid a growing globalpopulation.

  • Research Article
  • Cite Count Icon 46
  • 10.1016/j.asoc.2021.107788
Sentiment-influenced trading system based on multimodal deep reinforcement learning
  • Aug 11, 2021
  • Applied Soft Computing
  • Yu-Fu Chen + 1 more

Sentiment-influenced trading system based on multimodal deep reinforcement learning

  • Research Article
  • Cite Count Icon 1075
  • 10.1109/msp.2017.2738401
Deep Multimodal Learning: A Survey on Recent Advances and Trends
  • Nov 1, 2017
  • IEEE Signal Processing Magazine
  • Dhanesh Ramachandram + 1 more

The success of deep learning has been a catalyst to solving increasingly complex machine-learning problems, which often involve multiple data modalities. We review recent advances in deep multimodal learning and highlight the state-of the art, as well as gaps and challenges in this active research field. We first classify deep multimodal learning architectures and then discuss methods to fuse learned multimodal representations in deep-learning architectures. We highlight two areas of research–regularization strategies and methods that learn or optimize multimodal fusion structures–as exciting areas for future work.

  • Conference Article
  • 10.1109/icerect56837.2022.10060086
Multi-Modal Colour Extraction Using Deep Learning Techniques
  • Dec 26, 2022
  • Karthik Kulkarni + 2 more

Multimodal learning research has advanced quickly over the past ten years in a variety of fields, especially computer vision. Due to the growing opportunities of multimodal streaming data and deep learning algorithms, deep multimodal learning is increasingly common. This calls for the development of models that can reliably handle and interpret the multimodal data. Unstructured real-world data, often known as modalities, can naturally assume many different shapes, including both text and images often. Deep learning researchers are still driven by the need to extract useful patterns from this type of data. It is crucial to have well-organized product catalogues to enhance customers' experiences as they explore the plethora of possibilities provided by online marketplaces. The availability of product characteristics like colour or material is a crucial component of it. However, attribute data is frequently erroneous or absent on several of the markets we focus on. Utilizing deep models that have been trained on huge corpora to predict features from unstructured data, such as product descriptions and photographs (referred to as modalities in this study), is one potential approach to solving this issue. To receive a comprehensive rundown of the various multi-modal colour extraction techniques and their advantages, drawbacks and open challenges.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 74
  • 10.1007/s13202-021-01087-4
Prediction performance advantages of deep machine learning algorithms for two-phase flow rates through wellhead chokes
  • Feb 23, 2021
  • Journal of Petroleum Exploration and Production
  • Hossein Shojaei Barjouei + 6 more

Two-phase flow rate estimation of liquid and gas flow through wellhead chokes is essential for determining and monitoring production performance from oil and gas reservoirs at specific well locations. Liquid flow rate (QL) tends to be nonlinearly related to these influencing variables, making empirical correlations unreliable for predictions applied to different reservoir conditions and favoring machine learning (ML) algorithms for that purpose. Recent advances in deep learning (DL) algorithms make them useful for predicting wellhead choke flow rates for large field datasets and suitable for wider application once trained. DL has not previously been applied to predict QL from a large oil field. In this study, 7245 multi-well data records from Sorush oil field are used to compare the QL prediction performance of traditional empirical, ML and DL algorithms based on four influencing variables: choke size (D64), wellhead pressure (Pwh), oil specific gravity (γo) and gas–liquid ratio (GLR). The prevailing flow regime for the wells evaluated is critical flow. The DL algorithm substantially outperforms the other algorithms considered in terms of QL prediction accuracy. The DL algorithm predicts QL for the testing subset with a root-mean-squared error (RMSE) of 196 STB/day and coefficient of determination (R2) of 0.9969 for Sorush dataset. The QL prediction accuracy of the models evaluated for this dataset can be arranged in the descending order: DL > DT > RF > ANN > SVR > Pilehvari > Baxendell > Ros > Glbert > Achong. Analysis reveals that input variable GLR has the greatest, whereas input variable D64 has the least relative influence on dependent variable QL.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 84
  • 10.1007/s13202-022-01531-z
Predicting shear wave velocity from conventional well logs with deep and hybrid machine learning algorithms
  • Jul 11, 2022
  • Journal of Petroleum Exploration and Production Technology
  • Meysam Rajabi + 9 more

Shear wave velocity (VS) data from sedimentary rock sequences is a prerequisite for implementing most mathematical models of petroleum engineering geomechanics. Extracting such data by analyzing finite reservoir rock cores is very costly and limited. The high cost of sonic dipole advanced wellbore logging service and its implementation in a few wells of a field has placed many limitations on geomechanical modeling. On the other hand, shear wave velocity VS tends to be nonlinearly related to many of its influencing variables, making empirical correlations unreliable for its prediction. Hybrid machine learning (HML) algorithms are well suited to improving predictions of such variables. Recent advances in deep learning (DL) algorithms suggest that they too should be useful for predicting VS for large gas and oil field datasets but this has yet to be verified. In this study, 6622 data records from two wells in the giant Iranian Marun oil field (MN#163 and MN#225) are used to train HML and DL algorithms. 2072 independent data records from another well (MN#179) are used to verify the VS prediction performance based on eight well-log-derived influencing variables. Input variables are standard full-set recorded parameters in conventional oil and gas well logging data available in most older wells. DL predicts VS for the supervised validation subset with a root mean squared error (RMSE) of 0.055 km/s and coefficient of determination (R2) of 0.9729. It achieves similar prediction accuracy when applied to an unseen dataset. By comparing the VS prediction performance results, it is apparent that the DL convolutional neural network model slightly outperforms the HML algorithms tested. Both DL and HLM models substantially outperform five commonly used empirical relationships for calculating VS from Vp relationships when applied to the Marun Field dataset. Concerns regarding the model's integrity and reproducibility were also addressed by evaluating it on data from another well in the field. The findings of this study can lead to the development of knowledge of production patterns and sustainability of oil reservoirs and the prevention of enormous damage related to geomechanics through a better understanding of wellbore instability and casing collapse problems.Graphical abstract

  • Research Article
  • Cite Count Icon 60
  • 10.1016/j.imavis.2025.105509
A systematic review of intermediate fusion in multimodal deep learning for biomedical applications
  • May 1, 2025
  • Image and Vision Computing
  • Valerio Guarrasi + 6 more

Deep learning has revolutionized biomedical research by providing sophisticated methods to handle complex, high-dimensional data. Multimodal deep learning (MDL) further enhances this capability by integrating diverse data types such as imaging, textual data, and genetic information, leading to more robust and accurate predictive models. In MDL, differently from early and late fusion methods, intermediate fusion stands out for its ability to effectively combine modality-specific features during the learning process. This systematic review comprehensively analyzes and formalizes current intermediate fusion methods in biomedical applications, highlighting their effectiveness in improving predictive performance and capturing complex inter-modal relationships. We investigate the techniques employed, the challenges faced, and potential future directions for advancing intermediate fusion methods. Additionally, we introduce a novel structured notation that standardizes intermediate fusion architectures, enhancing understanding and facilitating implementation across various domains. Our findings provide actionable insights and practical guidelines intended to support researchers, healthcare professionals, and the broader deep learning community in developing more sophisticated and insightful multimodal models. Through this review, we aim to provide a foundational framework for future research and practical applications in the dynamic field of MDL. • Comprehensive review of intermediate fusion in multimodal learning in biomedicine. • Structured notation for categorizing intermediate fusion methods. • Analysis of the benefits and challenges of intermediate fusion in biomedical contexts. • Identification of future research directions for improving current fusion techniques. • Versatile framework applicable to other multimodal deep learning domains.

  • Research Article
  • Cite Count Icon 163
  • 10.1145/3545572
A Review on Methods and Applications in Multimodal Deep Learning
  • Feb 17, 2023
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Summaira Jabeen + 5 more

Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning (MMDL) is to create models that can process and link information using various modalities. Despite the extensive development made for unimodal learning, it still cannot cover all the aspects of human learning. Multimodal learning helps to understand and analyze better when various senses are engaged in the processing of information. This article focuses on multiple types of modalities, i.e., image, video, text, audio, body gestures, facial expressions, physiological signals, flow, RGB, pose, depth, mesh, and point cloud. Detailed analysis of the baseline approaches and an in-depth study of recent advancements during the past five years (2017 to 2021) in multimodal deep learning applications has been provided. A fine-grained taxonomy of various multimodal deep learning methods is proposed, elaborating on different applications in more depth. Last, main issues are highlighted separately for each domain, along with their possible future research directions.

  • PDF Download Icon
  • Supplementary Content
  • Cite Count Icon 19
  • 10.3390/e26030235
Deep Learning for 3D Reconstruction, Augmentation, and Registration: A Review Paper
  • Mar 7, 2024
  • Entropy
  • Prasoon Kumar Vinodkumar + 4 more

The research groups in computer vision, graphics, and machine learning have dedicated a substantial amount of attention to the areas of 3D object reconstruction, augmentation, and registration. Deep learning is the predominant method used in artificial intelligence for addressing computer vision challenges. However, deep learning on three-dimensional data presents distinct obstacles and is now in its nascent phase. There have been significant advancements in deep learning specifically for three-dimensional data, offering a range of ways to address these issues. This study offers a comprehensive examination of the latest advancements in deep learning methodologies. We examine many benchmark models for the tasks of 3D object registration, augmentation, and reconstruction. We thoroughly analyse their architectures, advantages, and constraints. In summary, this report provides a comprehensive overview of recent advancements in three-dimensional deep learning and highlights unresolved research areas that will need to be addressed in the future.

  • Book Chapter
  • Cite Count Icon 39
  • 10.1007/978-3-030-17795-9_5
Nature Inspired Meta-heuristic Algorithms for Deep Learning: Recent Progress and Novel Perspective
  • Apr 24, 2019
  • Haruna Chiroma + 6 more

Deep learning is presently attracting extra ordinary attention from both the industry and the academia. The application of deep learning in computer vision has recently gain popularity. The optimization of deep learning models through nature inspired algorithms is a subject of debate in computer science. The application areas of the hybrid of natured inspired algorithms and deep learning architecture includes: machine vision and learning, image processing, data science, autonomous vehicles, medical image analysis, biometrics, etc. In this paper, we present recent progress on the application of nature inspired algorithms in deep learning. The survey pointed out recent development issues, strengths, weaknesses and prospects for future research. A new taxonomy is created based on natured inspired algorithms for deep learning. The trend of the publications in this domain is depicted; it shows the research area is growing but slowly. The deep learning architectures not exploit by the nature inspired algorithms for optimization are unveiled. We believed that the survey can facilitate synergy between the nature inspired algorithms and deep learning research communities. As such, massive attention can be expected in a near future.

  • Research Article
  • Cite Count Icon 4
  • 10.1016/j.jdent.2023.104588
Multi-modal deep learning for automated assembly of periapical radiographs
  • Jun 21, 2023
  • Journal of Dentistry
  • L Pfänder + 5 more

Multi-modal deep learning for automated assembly of periapical radiographs

  • Research Article
  • Cite Count Icon 54
  • 10.1088/1361-6560/ac4c47
Deep multimodal learning for lymph node metastasis prediction of primary thyroid cancer
  • Feb 1, 2022
  • Physics in Medicine & Biology
  • Xinglong Wu + 3 more

Objective. The incidence of primary thyroid cancer has risen steadily over the past decades because of overdiagnosis and overtreatment through the improvement in imaging techniques for screening, especially in ultrasound examination. Metastatic status of lymph nodes is important for staging the type of primary thyroid cancer. Deep learning algorithms based on ultrasound images were thus developed to assist radiologists on the diagnosis of lymph node metastasis. The objective of this study is to integrate more clinical context (e.g., health records and various image modalities) into, and explore more interpretable patterns discovered by, deep learning algorithms for the prediction of lymph node metastasis in primary thyroid cancer patients. Approach. A deep multimodal learning network was developed in this study with a novel index proposed to compare the contribution of different modalities when making the predictions. Main results. The proposed multimodal network achieved an average F1 score of 0.888 and an average area under the receiver operating characteristic curve (AUC) value of 0.973 in two independent validation sets, and the performance was significantly better than that of three single-modality deep learning networks. Moreover, among three modalities used in this study, the deep multimodal learning network relied generally more on image modalities than the data modality of clinic records when making the predictions. Significance. Our work is beneficial to prospective clinic trials of radiologists on the diagnosis of lymph node metastasis in primary thyroid cancer, and will better help them understand how the predictions are made in deep multimodal learning algorithms.

  • Dissertation
  • 10.32657/10356/182346
Data efficient deep multimodal learning
  • Jan 1, 2025
  • Meng Shen

Multimodal learning, which enables neural networks to process and integrate information from various sensory modalities such as vision, language, and sound, has become increasingly important in applications ranging from affective computing and healthcare to advanced multimodal chatbots. Despite its potential, multimodal learning faces significant challenges, particularly in the area of data efficiency. The requirement for large, high-quality datasets from multiple modalities presents a substantial barrier, limiting the scalability and accessibility of large multimodal models. This dissertation addresses several key issues in data-efficient deep multimodal learning, focusing on the imbalanced multimodal data selection, the cold-start problem in multimodal active learning, and the mitigation of hallucinations in large vision-language models. Firstly, we analyze the limitations of conventional active learning strategies, which tend to favor dominant modalities, leading to unbalanced multimodal models that neglect weaker modalities. To overcome this, we propose a gradient embedding modulation method that ensures a more equitable data selection process across modalities, resulting in models that fairly uilize both weak and strong modalities. Building on our work in warm-start active learning, we tackle the cold-start problem in multimodal active learning, where no initial labels are available for warm-start data selection. We develop a two-stage approach that first reduces the modality representation gap through multimodal self-supervised learning, utilizing unimodal prototypes to harmonize representations across modalities. In the subsequent data selection stage, we introduce a regularization term to maximize modality alignment, leading to improved model performance using the same amount of data compared to existing methods. Extending our focus from data selection to the usage of training data, we address the challenge of hallucinations in large vision-language models, where the models generate content that is incorrect in the context of input images. We investigate the relationship between hallucinations and visual dependence of tokens, revealing that certain tokens contribute disproportionately to these hallucinatory. Based on this insight, we propose an approach that adjusts training weights according to the visual dependence of tokens, thereby reducing the hallucination rate without requiring additional training data or inference costs. The contributions of this thesis offer significant advancements in the field of dataefficient multimodal learning. By developing novel methods for balancing multimodal data selection, addressing cold-start problem in multimodal active learning, and mitigating hallucinations in large vision-language models, this work paves the way for more practical and scalable multimodal learning systems that require less data and computational effort while achieving superior performance.

  • Research Article
  • Cite Count Icon 3
  • 10.57238/n65d0p57
Deep Learning for Computer Vision: Innovations in Image Recognition and Processing Techniques
  • Jun 30, 2024
  • CyberSystem Journal
  • Akeel Mahmoud + 1 more

Deep learning is a key area of research in the field of computer vision, image processing and bioinformatics. The techniques of deep learning generally are divided into three categories namely Convolutional Neural Networks (CNN), Restricted Boltzmann Machines (RBM), Stacked RBM and HOG (Histograms of oriented Gradient) feature extraction, Convolutional Neural Networks as a Database (CNN as D). Additionally, one in few deep learning architectures which is gaining popularity and is frequently used in the field of computer vision and image processing is Extreme Learning Machine and ensemble of Extreme Learning Machine and CNN. It attempts to survey the recent advances in deep learning researchers and the application of these algorithm in the field of computer vision. Mainly focusing on the deep learning methods and algorithms rather than image processing and computer vision methods, this work inspects deep learning techniques which are widely and commonly used in the field of computer vision image detection and processing like CNN, DBN, RBM and HMM as well as various applications of these techniques. Applications of deep learning techniques in computer vision are image classification, object recognition and detection. Along with the recent works and the future scope for deep learning methods in the field of computer vision and image processing is presented.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant