Articles published on Image fusion
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
14582 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.net.2026.104294
- Jul 1, 2026
- Nuclear Engineering and Technology
- Shuo Li + 7 more
Study on image fusion methods for spatial distribution localization of radiation sources
- New
- Research Article
- 10.1016/j.soc.2025.12.003
- Jul 1, 2026
- Surgical oncology clinics of North America
- Mariano E Giménez + 4 more
Future of Robotics and Integration of Artificial Intelligence: Toward Computer-Assisted Surgery and the Real Democratization of Surgical Care.
- New
- Research Article
- 10.1016/j.inffus.2026.104171
- Jul 1, 2026
- Information Fusion
- Yiheng Wang + 3 more
Multimodal fusion of 3D point cloud and intraoperative imaging to enhance surgical robot navigation
- New
- Research Article
- 10.1016/j.inffus.2026.104149
- Jul 1, 2026
- Information Fusion
- Weixuan Ma + 6 more
GeoCraft: A Diffusion Model-based 3D Reconstruction Method driven by image and point cloud fusion
- New
- Research Article
- 10.1016/j.image.2026.117549
- Jul 1, 2026
- Signal Processing: Image Communication
- Junhao He + 3 more
MDCFusion: Enhancing infrared and visible image fusion through Multi-Scale Dense Convolutional Sparse Coding
- New
- Research Article
- 10.1016/j.patcog.2026.113130
- Jul 1, 2026
- Pattern Recognition
- Wei Tang + 3 more
SETFusion: A semantic transformer for infrared and visible image fusion
- New
- Research Article
- 10.1016/j.dsp.2026.106109
- Jul 1, 2026
- Digital Signal Processing
- Phu-Hung Dinh + 3 more
HDAFusion: Hierarchical decomposition and attention-based framework for infrared and visible image fusion
- New
- Research Article
- 10.1016/j.engappai.2026.114754
- Jul 1, 2026
- Engineering Applications of Artificial Intelligence
- Xiaowen Liu + 4 more
An anisotropic low-rank self-attention mechanism for infrared and visible image fusion
- New
- Research Article
- 10.1080/17538947.2026.2639814
- Jul 1, 2026
- International Journal of Digital Earth
- Mengnan Jin + 4 more
ABSTRACT Hyperspectral images generally suffer from low spatial resolution, limiting their utility in fine-scale applications. To address the limitations of existing unsupervised fusion methods in adapting to spatial heterogeneity and scale effects that often lead to blurred details and reduced target discrimination, we propose an unsupervised adaptive scale-aware detail feature extraction network (UASNet), which introduces a structure-adaptive mechanism and degradation-aware modeling to effectively harmonize the representation consistency of multiscale ground objects without paired training samples. The network consists of three stages: the prior information mining stage, the spectral channel mapping stage, and the detail feature fusion stage. Specifically, an adaptive scale-aware convolution is embedded within a reversible detail extraction module to capture features of objects with varying scales and geometries, thereby preserving fine textures and structural integrity. Furthermore, driven by the requirements of real-world application scenarios, spectral and spatial branches are constructed to learn the respective degradation priors, which are integrated into the loss function to realize a fully unsupervised fusion framework. Extensive experiments on both simulated and real datasets demonstrate the superior performance of the proposed method compared with that of state-of-the-art approaches. Moreover, even without ground truth data, the downstream classification results indirectly validate that UASNet exhibits strong potential for real-world applications.
- New
- Research Article
3
- 10.1016/j.inffus.2026.104130
- Jul 1, 2026
- Information Fusion
- Xilai Li + 5 more
All-weather multi-modality image fusion: Unified framework and 100k benchmark
- New
- Research Article
- 10.1016/j.ultramic.2026.114388
- Jul 1, 2026
- Ultramicroscopy
- Ethan Saul Carrizales Alvarez + 2 more
Component substitution and multiresolution analysis on hyperspectral microwave microscopy data sets.
- New
- Research Article
- 10.1016/j.bspc.2026.110043
- Jul 1, 2026
- Biomedical Signal Processing and Control
- Yalin Li + 4 more
Semantic–geometric dual alignment: A progressive co-optimization paradigm for misaligned multimodal medical image fusion
- New
- Research Article
- 10.1016/j.image.2026.117552
- Jul 1, 2026
- Signal Processing: Image Communication
- Yuyuan Luo + 1 more
Hyperspectral and multispectral image fusion via N-gram transformer and local adaptive fusion strategy
- New
- Research Article
- 10.1016/j.bspc.2026.110075
- Jul 1, 2026
- Biomedical Signal Processing and Control
- Hegui Zhu + 3 more
A multi-scale expert gating network with cross-modal attention for multimodal medical image fusion
- New
- Research Article
- 10.1038/s41598-026-57893-5
- Jun 30, 2026
- Scientific reports
- Marcel Brettmacher + 5 more
Histological analysis traditionally relies on thin tissue sections, providing inherently two-dimensional (2D) information. However, this approach captures only a fraction of the entire sample and lacks the spatial context necessary for comprehensive tissue assessment. Recent advancements in multimodal imaging have introduced the fusion of histological data with three-dimensional (3D) imaging techniques, such as Light Sheet Fluorescence Microscopy (LSFM), to enhance tissue analysis by integrating complementary spatial information. A key challenge in this fusion process is the accurate alignment of corresponding structures across modalities, which is complicated by differences in resolution, sectioning-induced deformations, and varying imaging orientations. This is further complicating in the case of 2D-to-3D registration where the initial alignment of the image inside the volume is unknown and registration processes are computationally expensive due to six degrees of freedom in the placement. Here, existing solutions often require manual selection of image pairs, fiducial markers or technical expertise, limiting accessibility to non-specialist users. To address these limitations, we introduce LitSHi (Light Sheet meets Histology), a novel registration tool that enables the automated and precise alignment of LSFM and histological images. LitSHi allows multimodal image fusion to be performed fully automatically, which significantly reduces the need for manual intervention. Using testicular tumor and pancreatic specimens, we evaluated LitSHi and demonstrated its ability to enhance structural correspondence between LSFM and histological images. The automated registration process markedly improved both efficiency and alignment accuracy compared with conventional manual or semi-automated approaches. Overall, LitSHi holds significant potential to advance digital pathology by enabling optimized multimodal tissue analysis and supporting future developments in computational pathology and AI-driven diagnostics.
- New
- Research Article
- 10.1038/s41598-026-59063-z
- Jun 28, 2026
- Scientific reports
- J Bineeshia + 1 more
Infrared-visible (IR-VI) image fusion attempts to combine complementary thermal saliency and rich visual detail into a single valuable representation. However, contemporary fusion algorithms frequently rely primarily on low-level feature similarity or modality-agnostic attention, which results in inadequate retention of semantic relevance and object-level importance especially in complex scenarios. This research suggests Context-Aware and Semantic-Enhanced Generative Adversarial Network (CASE-GANet) that explicitly integrates semantic knowledge into the fusion process in order to overcome these constraints. Modality-specific feature extractors are used, such as an attentional MobileNet-V2 with Coordinate Attention for visible image data to capture texture-rich spatial information and a Residual Dense Block (RDB) network for infrared images to preserve weak thermal structures. A Context Score Generator uses the high-level scene comprehension produced by a semantic segmentation branch to evaluate class-wise modality relevance. Semantically guided and spatially adaptive cross-modal feature fusion is made possible by injecting this contextual information into a unique Context-Aware Multi-Head Attention (C-MHA) module via a learnable bias map. In both qualitative and quantitative assessments, extensive experiments on the TNO and Road Scene, show that the proposed strategy outperforms the latest fusion techniques. The findings verify that integrating semantic context to attention-based fusion improves overall perceptual quality, structural fidelity, and thermal target visibility.
- New
- Research Article
- 10.1021/acs.analchem.6c00810
- Jun 26, 2026
- Analytical chemistry
- Wenjing Chen + 3 more
Prostate cancer (PCa) is one of the most common malignant tumors in men, necessitating the use of effective methods for early detection. This study proposes a multimodal approach based on urine analysis using laser-induced breakdown spectroscopy (LIBS) and Fourier transform infrared spectroscopy (FTIR). Specifically, we developed a dual-spectrum reconstructed image fusion network (SMFNet) incorporating a spectrum-to-image reconstruction strategy and optimized loss functions (feature margin and focus loss). By leveraging the complementarity between atomic and molecular spectral data, the SMFNet model enhances feature characterization and interclass differentiation. The results demonstrate that SMFNet achieves an accuracy of 97.62% and a macro-F1 of 98.85% on the test set, significantly outperforming single-modal methods and baseline models. Consequently, this method offers a novel, efficient, and reliable approach for the future early detection of PCa, with the potential to enhance diagnostic accuracy, shorten detection cycles, and provide a supportive basis for early clinical intervention.
- New
- Research Article
- 10.7507/1001-5515.202411017
- Jun 25, 2026
- Sheng wu yi xue gong cheng xue za zhi = Journal of biomedical engineering = Shengwu yixue gongchengxue zazhi
- Xuelian Gu + 6 more
Cervical intraepithelial neoplasia is the primary type of cervical precancerous lesion; however, manual clinical diagnosis is prone to bias and has limited grading accuracy. To achieve precise automated grading of CIN, this paper proposes a multimodal fusion Swin Transformer model and develops a corresponding computer-aided diagnosis system. This method employs three-channel fusion of raw images, cervical mask images, and directional gradient histogram features to enhance lesion texture and location information. Within the Swin Transformer backbone, an atrous spatial pyramid pooling module channel attention module and a convolutional feature extraction module are embedded to balance global semantic and local detail features. A focal loss function is adopted to address class imbalance in the dataset and improve the model's ability to identify difficult-to-classify samples. On a dataset of 3 915 clinical colposcopy images, the model achieved an overall accuracy of 90.01%, precision of 87.55%, recall of 86.17%, F1 score of 89.13%, outperforming baseline models such as VGG, ResNet, and Swin Transformer. The developed system integrates image quality screening, lesion identification, and three-level classification functions, providing an effective tool for the rapid and objective screening of clinical cervical precancerous lesions.
- New
- Research Article
- 10.1109/tip.2026.3703760
- Jun 22, 2026
- IEEE transactions on image processing : a publication of the IEEE Signal Processing Society
- Andong Lu + 5 more
Pixel-level fusion is widely considered a lightweight yet limited strategy in RGB-Thermal (RGBT) tracking due to its shallow representational capacity. However, its actual limitations and potential remain largely unexplored. We systematically analyze fusion location, modality alignment, and tracking performance, revealing that despite lower modality gaps than feature-level fusion, pixel-level fusion lacks task-relevant discrimination, restricting its effectiveness. In this paper, we propose the Task-driven Pixel-level Fusion tracker (TPF), which preserves the efficiency of early fusion while enhancing discriminative capacity. Central to TPF is a lightweight pixel fusion adapter that ensures real-time image fusion with only 14.3KB extra parameters over the baseline at inference. To enhance its limited representational capacity, we propose a task-driven progressive learning framework consisting of two key stages. First, a heterogeneous multi-expert distillation scheme adaptively transfers image fusion knowledge from diverse models under tracking-guided evaluation, mitigating the generalization limitations of single-teacher distillation across varied tracking scenarios. Second, to overcome limited task discrimination caused by sparse, target-focused tracking supervision, we propose a de coupled representation learning strategy that offers dense, complementary guidance to improve target-background separation and fusion quality. A nearest-neighbor dynamic template update further enhances robustness to appearance changes. Extensive experiments on four RGBT tracking benchmarks show that TPF achieves competitive accuracy and speed, outperforming both feature-level and existing pixel-level fusion methods, offering new insights into efficient RGBT tracking.
- New
- Research Article
- 10.1038/s41598-026-58323-2
- Jun 21, 2026
- Scientific reports
- Guo Chen + 4 more
Based on the complementary and enhanced fusion of 3D point clouds and 2D RGB images, this paper designs an end-to-end learning framework-Point Cloud Enhanced Depth Pixel Fusion Network (PEPF-Net), aimed at enabling robots to achieve accurate 3D perception of unstructured environments. In the process, we address four key problems in 3D perception tasks: enhancing RGB representation using the reflection intensity and depth information of point clouds to generate Depth-RGB Pixel (D-Pixel); proposing Point-by-Point Vector Attention (PVA-Net) to model the vector relationships of point clouds, to obtain deep-level point cloud features, and to achieve direct and effective fusion of heterogeneous data; designing a Layered-Transformer (L-TsfmNet) feature extractor to hierarchically extract D-Pixel features; proposing Variable Window Self-attention (VS-a) to focus on the relationships between local "window tokens" and avoid the complexity of global computation. Extensive experiments on the KITTI dataset demonstrate that PEPF-Net outperforms the currently common advanced environmental 3D perception algorithms.