Infrared and visible image fusion methods and applications: A survey
Infrared and visible image fusion methods and applications: A survey
- Research Article
8
- 10.3788/irla20200467
- Jan 1, 2021
- Infrared and Laser Engineering
Infrared and visible image fusion combines the infrared thermal radiation information and visible detail information. The image fusion technique has facilitated development in numerous fields, including production, life sciences, military surveillance and others, and has become a key research direction in the field of image technology. According to the core idea, fusion framework and research progress of image fusion methods, the fusion methods based on multi-scale transformation, sparse representation, neural network, etc. are elaborated and compared, and the application status of infrared and visible light image fusion in various fields and the commonly used the evaluation index. The most representative methods and evaluation indicators are selected and applied to six different scenes in order to verify the advantages and disadvantages of each one. Finally, the existing problems of infrared and visible image fusion methods are experimentally analyzed and summarized , the development prospects of infrared and visible image fusion technology are presented.
- Research Article
- 10.1371/journal.pone.0318931
- Mar 28, 2025
- PloS one
In response to the limitations of current infrared and visible light image fusion algorithms-namely insufficient feature extraction, loss of detailed texture information, underutilization of differential and shared information, and the high number of model parameters-this paper proposes a novel multi-scale infrared and visible image fusion method with two-branch feature interaction. The proposed method introduces a lightweight multi-scale group convolution, based on GS convolution, which enhances multi-scale information interaction while reducing network parameters by incorporating group convolution and stacking multiple small convolutional kernels. Furthermore, the multi-level attention module is improved by integrating edge-enhanced branches and depthwise separable convolutions to preserve detailed texture information. Additionally, a lightweight cross-attention fusion module is introduced, optimizing the use of differential and shared features while minimizing computational complexity. Lastly, the efficiency of local attention is enhanced by adding a multi-dimensional fusion branch, which bolsters the interaction of information across multiple dimensions and facilitates comprehensive spatial information extraction from multimodal images. The proposed algorithm, along with seven others, was tested extensively on public datasets such as TNO and Roadscene. The experimental results demonstrate that the proposed method outperforms other algorithms in both subjective and objective evaluation results. Additionally, it demonstrates good performance in terms of operational efficiency. Moreover, target detection performance experiments conducted on the [Formula: see text] dataset confirm the superior performance of the proposed algorithm.
- Research Article
5
- 10.1016/j.neucom.2024.128957
- Nov 20, 2024
- Neurocomputing
DSAFuse: Infrared and visible image fusion via dual-branch spatial adaptive feature extraction
- Research Article
16
- 10.3390/s22176390
- Aug 25, 2022
- Sensors (Basel, Switzerland)
The purpose of infrared and visible image fusion is to generate images with prominent targets and rich information which provides the basis for target detection and recognition. Among the existing image fusion methods, the traditional method is easy to produce artifacts, and the information of the visible target and texture details are not fully preserved, especially for the image fusion under dark scenes and smoke conditions. Therefore, an infrared and visible image fusion method is proposed based on visual saliency image and image contrast enhancement processing. Aiming at the problem that low image contrast brings difficulty to fusion, an improved gamma correction and local mean method is used to enhance the input image contrast. To suppress artifacts that are prone to occur in the process of image fusion, a differential rolling guidance filter (DRGF) method is adopted to decompose the input image into the basic layer and the detail layer. Compared with the traditional multi-scale decomposition method, this method can retain specific edge information and reduce the occurrence of artifacts. In order to solve the problem that the salient object of the fused image is not prominent and the texture detail information is not fully preserved, the salient map extraction method is used to extract the infrared image salient map to guide the fusion image target weight, and on the other hand, it is used to control the fusion weight of the basic layer to improve the shortcomings of the traditional ‘average’ fusion method to weaken the contrast information. In addition, a method based on pixel intensity and gradient is proposed to fuse the detail layer and retain the edge and detail information to the greatest extent. Experimental results show that the proposed method is superior to other fusion algorithms in both subjective and objective aspects.
- Research Article
29
- 10.1007/s40747-022-00722-9
- Apr 22, 2022
- Complex & Intelligent Systems
For the past few years, image fusion technology has made great progress, especially in infrared and visible light image infusion. However, the fusion methods, based on traditional or deep learning technology, have some disadvantages such as unobvious structure or texture detail loss. In this regard, a novel generative adversarial network named MSAt-GAN is proposed in this paper. It is based on multi-scale feature transfer and deep attention mechanism feature fusion, and used for infrared and visible image fusion. First, this paper employs three different receptive fields to extract the multi-scale and multi-level deep features of multi-modality images in three channels rather than artificially setting a single receptive field. In this way, the important features of the source image can be better obtained from different receptive fields and angles, and the extracted feature representation is also more flexible and diverse. Second, a multi-scale deep attention fusion mechanism is designed in this essay. It describes the important representation of multi-level receptive field extraction features through both spatial and channel attention and merges them according to the level of attention. Doing so can lay more emphasis on the attention feature map and extract significant features of multi-modality images, which eliminates noise to some extent. Third, the concatenate operation of the multi-level deep features in the encoder and the deep features in the decoder are cascaded to enhance the feature transmission while making better use of the previous features. Finally, this paper adopts a dual-discriminator generative adversarial network on the network structure, which can force the generated image to retain the intensity of the infrared image and the texture detail information of the visible image at the same time. Substantial qualitative and quantitative experimental analysis of infrared and visible image pairs on three public datasets show that compared with state-of-the-art fusion methods, the proposed MSAt-GAN network has comparable outstanding fusion performance in subjective perception and objective quantitative measurement.
- Research Article
1
- 10.1117/1.jei.29.2.023016
- Apr 1, 2020
- Journal of Electronic Imaging
The infrared image (RI) and visible image (VI) fusion method merges complementary information from the infrared and visible imaging sensors to provide an effective way for understanding the scene. The graph filter bank-based graph wavelet transform possesses the advantages of the classic wavelet filter bank and graph representation of a signal. Therefore, we propose an RI and VI fusion method based on oversampled graph filter banks. Specifically, we consider the source images as signals on the regular graph and decompose them into the multiscale representations with M-channel oversampled graph filter banks. Then, the fusion rule for the low-frequency subband is constructed using the modified local coefficient of variation and the bilateral filter. The fusion maps of detail subbands are formed using the standard deviation-based local properties. Finally, the fusion image is obtained by applying the inverse transform on the fusion subband coefficients. The experimental results on benchmark images show the potential of the proposed method in the image fusion applications.
- Research Article
20
- 10.1109/access.2021.3090436
- Jan 1, 2021
- IEEE Access
The fusion quality of infrared and visible image is very important for subsequent human understanding of image information and target processing. The fusion quality of the existing infrared and visible image fusion methods still has room for improvement in terms of image contrast, sharpness and richness of detailed information. To obtain better fusion performance, an infrared and visible image fusion algorithm based on latent low-rank representation (LatLRR) nested with rolling guided image filtering (RGIF) is proposed that is a novel solution that integrates two-level decomposition and three-layer fusion. First, infrared and visible images are decomposed using LatLRR to obtain the low-rank sublayers, saliency sublayers, and sparse noise sublayers. Then, RGIF is used to perform further multiscale decomposition of the low-rank sublayers to extract multiple detail layers, which are fused using convolutional neural network (CNN)-based fusion rules to obtain the detail-enhanced layer. Next, an algorithm based on improved visual saliency mapping with weighted guided image filtering (IVSM-GIF) is used to fuse the low-rank sublayers, and an algorithm for adaptive weighting of regional energy features based on Laplacian pyramid decomposition is used to fuse the saliency sublayers. Finally, the fused low-rank sublayer, saliency sublayer, and detail-enhanced layer are used to reconstruct the final image. The experimental results show that the proposed method outperforms other state-of-the-art fusion methods in terms of visual quality and objective evaluation, achieving the highest average values in six objective evaluation metrics.
- Research Article
9
- 10.3390/e24020294
- Feb 19, 2022
- Entropy
In this paper, we design an infrared (IR) and visible (VIS) image fusion via unsupervised dense networks, termed as TPFusion. Activity level measurements and fusion rules are indispensable parts of conventional image fusion methods. However, designing an appropriate fusion process is time-consuming and complicated. In recent years, deep learning-based methods are proposed to handle this problem. However, for multi-modality image fusion, using the same network cannot extract effective feature maps from source images that are obtained by different image sensors. In TPFusion, we can avoid this issue. At first, we extract the textural information of the source images. Then two densely connected networks are trained to fuse textural information and source image, respectively. By this way, we can preserve more textural details in the fused image. Moreover, loss functions we designed to constrain two densely connected convolutional networks are according to the characteristics of textural information and source images. Through our method, the fused image will obtain more textural information of source images. For proving the validity of our method, we implement comparison and ablation experiments from the qualitative and quantitative assessments. The ablation experiments prove the effectiveness of TPFusion. Being compared to existing advanced IR and VIS image fusion methods, our fusion results possess better fusion results in both objective and subjective aspects. To be specific, in qualitative comparisons, our fusion results have better contrast ratio and abundant textural details. In quantitative comparisons, TPFusion outperforms existing representative fusion methods.
- Research Article
14
- 10.1109/access.2021.3056888
- Jan 1, 2021
- IEEE Access
Image fusion is a visual enhancement technique that combines source images from different sensors to produce a more robust and informative fused image for subsequent processing or decision making. Infrared and visible light images share complementary properties that enable the production of robust and informative fused images. This paper proposed an infrared and visible image fusion method that improved the tetrolet framework to improve infrared and visible image fusion quality. First, the source image is enhanced by bicubic interpolation. The improved tetrolet transform then decomposes the enhanced source image; the high-frequency components are fused by convolutional sparse representation theory and combined with corresponding rules, and the low-frequency components are fused by defining ISER descriptors. Finally, we use the inverse transform to reconstruct the fused image. Qualitative and quantitative experimental results on five groups of typical infrared and visible image datasets demonstrate the proposed method's effectiveness. The proposed method exhibits better performances on subjective vision and objective indexes compared with the other state-of-the-art methods.
- Research Article
383
- 10.1109/tmm.2020.2997127
- Jun 5, 2020
- IEEE Transactions on Multimedia
Infrared and visible image fusion aims to describe the same scene from different aspects by combining complementary information of multi-modality images. The existing Generative adversarial networks (GAN) based infrared and visible image fusion methods cannot perceive the most discriminative regions, and hence fail to highlight the typical parts existing in infrared and visible images. To this end, we integrate multi-scale attention mechanism into both generator and discriminator of GAN to fuse infrared and visible images (AttentionFGAN). The multi-scale attention mechanism aims to not only capture comprehensive spatial information to help generator focus on the foreground target information of infrared image and background detail information of visible image, but also constrain the discriminators focus more on the attention regions rather than the whole input image. The generator of AttentionFGAN consists of two multi-scale attention networks and an image fusion network. Two multi-scale attention networks capture the attention maps of infrared and visible images respectively, so that the fusion network can reconstruct the fused image by paying more attention to the typical regions of source images. Besides, two discriminators are adopted to force the fused result keep more intensity and texture information from infrared and visible image respectively. Moreover, to keep more information of attention region from source images, an attention loss function is designed. Finally, the ablation experiments illustrate the effectiveness of the key parts of our method, and extensive qualitative and quantitative experiments on three public datasets demonstrate the advantages and effectiveness of AttentionFGAN compared with the other state-of-the-art methods.
- Conference Article
1
- 10.1145/3671151.3671306
- Apr 26, 2024
The purpose of image fusion is to extract the most useful features from source images of different modalities and integrate them into a single image, facilitating subsequent task processing. Therefore, image fusion is a hot research field. In multi-modal image fusion, the fusion task of infrared images and visible images has many advantages. Infrared images can utilize radiation differences to distinguish between targets and backgrounds, and they can work well under both day and night conditions. In contrast, visible images are more consistent with human visual habits and have clear texture details. Therefore, it is necessary to fuse these two types of images together, which can combine the advantages of thermal radiation information in infrared images and detailed texture information in visible images. This paper reviews the latest research progress and application areas of image fusion technology. Firstly, the basic concept of image fusion is introduced, and the importance and application prospects of image fusion in multiple fields are outlined. Then, the basic principles and commonly used methods of image fusion are discussed in detail. A summary and comparison of performance evaluation methods for image fusion technology are provided, and the future development trends of image fusion technology are prospected, emphasizing research directions such as image fusion in complex environments under deep learning and real-time system design.
- Research Article
279
- 10.1016/j.infrared.2017.07.010
- Jul 12, 2017
- Infrared Physics & Technology
A survey of infrared and visual image fusion methods
- Research Article
419
- 10.1016/j.inffus.2022.10.034
- Nov 5, 2022
- Information Fusion
DIVFusion: Darkness-free infrared and visible image fusion
- Conference Article
18
- 10.1117/12.2551830
- Apr 1, 2020
With the development of sensor technologies, imaging technology is developing more rapidly. What followed was the widespread use of image processing technology in many kinds of applications. For instance, image processing technology has been widely used in video surveillance, medical diagnosis, remote sensing detection and object tracking. As a sub-field of image processing technology, image fusion is the one of most studied technology. The aim of image fusion is to acquire an integrated image that contains more information. This integrated image is more conductive for a human or a machine to understand and mine the information contained in the image. In all kinds of image fusion, infrared (IR) and visible (VIS) image fusion is one of the most valuable multisource image fusion. When imaging the same scene using both IR and VIS imaging system, more information can be obtained, but more redundant information is generated. The IR sensor acquires the thermal radiation information of the object in a scene, so the object can also be detected when the lighting conditions are poor. The image acquired by VIS light sensors has more spectral information, clearer texture details, and higher spatial resolution. Thus, the scene can be described more completely by integrating the IR and VIS images into one image. Meanwhile, the scene can be readily understood by observers, and the information of the scene can be easily perceived. In this paper, an effective IR and VIS image fusion via non-subsampled shearlet transform (NSST) and pulse-coupled neural network (PCNN) in multi-scale morphological gradient (MSMG) domain is proposed. First, low frequency sub-image and high frequency sub-images are obtained through NSST. Then, the low frequency sub-image and high frequency sub-images are fused via a MSMG domain PCNN (MSMG-PCNN) strategy. Finally, the fused image is reconstructed by inverse NSST. Experimental results demonstrate that the proposed MSMG-PCNN-NSST algorithm performs effectively in most cases by qualitative and quantitative evaluation.
- Research Article
88
- 10.1016/j.inffus.2021.06.002
- Jun 12, 2021
- Information Fusion
Self-supervised feature adaption for infrared and visible image fusion