Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Dual-encoder Multi-Scale Refinement network for robust crack segmentation across diverse domains

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Dual-encoder Multi-Scale Refinement network for robust crack segmentation across diverse domains

Similar Papers
  • Research Article
  • Cite Count Icon 12
  • 10.1016/j.ejmp.2023.102595
Robust and efficient abdominal CT segmentation using shape constrained multi-scale attention network
  • May 11, 2023
  • Physica Medica
  • Nuo Tong + 4 more

Robust and efficient abdominal CT segmentation using shape constrained multi-scale attention network

  • Research Article
  • 10.1109/access.2026.3678115
LMA-Net: Light-guided Multi-scale Attention for Robust Soil Crack Segmentation under Complex Illumination
  • Jan 1, 2026
  • IEEE Access
  • Guang-Zhu Zhang + 3 more

Soil is highly susceptible to cracking under climatic factors such as rainfall and temperature, which degrades mechanical properties and durability and threatens the safety and service life of road infrastructure. Efficient and robust crack segmentation is therefore essential for infrastructure life-cycle performance. In real-world scenes, complex illumination conditions such as non-uniform lighting and shadows can substantially degrade segmentation accuracy. To address this problem, a light-guided multi-scale attention segmentation network (LMA-Net) is proposed for soil crack segmentation under complex illumination. The proposed model takes soil crack images captured under complex illumination conditions as input and outputs pixel-level crack segmentation maps. The network adopts a U-shaped encoder–decoder architecture and integrates an illumination-guided attention (IGA) module to suppress illumination interference, a multi-scale feature fusion (MSF) module to enhance crack representation across varying widths, and SE-ResBlocks to improve channel discriminability in deeper layers. In addition, a soil crack dataset was constructed under wetting–drying cycles and diverse complex illumination scenarios. A total of 164 raw images were acquired and preprocessed into 164 full-image inputs for model development, from which 1620 crack-containing patches were generated for training and validation. Experimental results show that LMA-Net achieves an IoU of 78.74% and an F1 score of 87.83% on the test set, outperforming the comparative segmentation networks by more than 4.64% in IoU and 3.07% in F1 score. Ablation studies further verify the effectiveness of the proposed modules, and qualitative evaluation on real-world engineering images demonstrates that LMA-Net has strong potential for soil crack segmentation under complex environmental conditions.

  • Research Article
  • Cite Count Icon 34
  • 10.1002/mp.13677
Endometrium segmentation on transvaginal ultrasound image using key-point discriminator.
  • Jul 31, 2019
  • Medical Physics
  • Hyenok Park + 7 more

Transvaginal ultrasound imaging provides useful information for diagnosing endometrial pathologies and reproductive health. Endometrium segmentation in transvaginal ultrasound (TVUS) images is very challenging due to ambiguous boundaries and heterogeneous textures. In this study, we developed a new segmentation framework which provides robust segmentation against ambiguous boundaries and heterogeneous textures of TVUS images. To achieve endometrium segmentation from TVUS images, we propose a new segmentation framework with a discriminator guided by four key points of the endometrium (namely, the endometrium cavity tip, the internal os of the cervix, and the two thickest points between the two basal layers on the anterior and posterior uterine walls). The key points of the endometrium are defined as meaningful points that are related to the characteristics of the endometrial morphology, namely the length and thickness of the endometrium. In the proposed segmentation framework, the key-point discriminator distinguishes a predicted segmentation map from a ground-truth segmentation map according to the key-point maps. Meanwhile, the endometrium segmentation network predicts accurate segmentation results that the key-point discriminator cannot discriminate. In this adversarial way, the key-point information containing endometrial morphology characteristics is effectively incorporated in the segmentation network. The segmentation network can accurately find the segmentation boundary while the key-point discriminator learns the shape distribution of the endometrium. Moreover, the endometrium segmentation can be robust to the heterogeneous texture of the endometrium. We conducted an experiment on a TVUS dataset that contained 3,372 sagittal TVUS images and the corresponding key points. The dataset was collected by three hospitals (Ewha Woman's University School of Medicine, Asan Medical Center, and Yonsei University College of Medicine) with the approval of the three hospitals' Institutional Review Board. For verification, fivefold cross-validation was performed. The proposed key-point discriminator improved the performance of the endometrium segmentation, achieving 82.67 % for the Dice coefficient and 70.46% for the Jaccard coefficient. In comparison, on the TVUS images UNet, showed 58.69 % for the Dice coefficient and 41.59 % for the Jaccard coefficient. The qualitative performance of the endometrium segmentation was also improved over the conventional deep learning segmentation networks. Our experimental results indicated robust segmentation by the proposed method on TVUS images with heterogeneous texture and unclear boundary. In addition, the effect of the key-point discriminator was verified by an ablation study. We proposed a key-point discriminator to train a segmentation network for robust segmentation of the endometrium with TVUS images. By utilizing the key-point information, the proposed method showed more reliable and accurate segmentation performance and outperformed the conventional segmentation networks both in qualitative and quantitative comparisons.

  • Research Article
  • Cite Count Icon 6
  • 10.1016/j.eswa.2024.124070
SCPMan: Shape context and prior constrained multi-scale attention network for pancreatic segmentation
  • May 2, 2024
  • Expert Systems With Applications
  • Leilei Zeng + 6 more

SCPMan: Shape context and prior constrained multi-scale attention network for pancreatic segmentation

  • Research Article
  • 10.1016/j.rineng.2026.110145
Hybrid multi-scale CNN-Transformer network for structural surface crack segmentation
  • Jun 1, 2026
  • Results in Engineering
  • Daniel Asefa Beyene + 5 more

Hybrid multi-scale CNN-Transformer network for structural surface crack segmentation

  • Research Article
  • Cite Count Icon 3
  • 10.32604/jihpp.2019.07198
A Multi-Scale Network with the Encoder-Decoder Structure for CMR Segmentation
  • Jan 1, 2019
  • Journal of Information Hiding and Privacy Protection
  • Chaoyang Xia + 3 more

Cardiomyopathy is one of the most serious public health threats. The precise structural and functional cardiac measurement is an essential step for clinical diagnosis and follow-up treatment planning. Cardiologists are often required to draw endocardial and epicardial contours of the left ventricle (LV) manually in routine clinical diagnosis or treatment planning period. This task is time-consuming and error-prone. Therefore, it is necessary to develop a fully automated end-to-end semantic segmentation method on cardiac magnetic resonance (CMR) imaging datasets. However, due to the low image quality and the deformation caused by heartbeat, there is no effective tool for fully automated end-to-end cardiac segmentation task. In this work, we propose a multi-scale segmentation network (MSSN) for left ventricle segmentation. It can effectively learn myocardium and blood pool structure representations from 2D short-axis CMR image slices in a multi-scale way. Specifically, our method employs both parallel and serial of dilated convolution layers with different dilation rates to capture multi-scale semantic features. Moreover, we design graduated up-sampling layers with subpixel layers as the decoder to reconstruct lost spatial information and produce accurate segmentation masks. We validated our method using 164 T1 Mapping CMR images and showed that it outperforms the advanced convolutional neural network (CNN) models. In validation metrics, we archived the Dice Similarity Coefficient (DSC) metric of 78.96%.

  • Research Article
  • Cite Count Icon 7
  • 10.1038/s41598-025-09351-x
AG-MS3D-CNN multiscale attention guided 3D convolutional neural network for robust brain tumor segmentation across MRI protocols
  • Jul 7, 2025
  • Scientific Reports
  • Umesh Kumar Lilhore + 7 more

Accurate segmentation of brain tumors from multimodal Magnetic Resonance Imaging (MRI) plays a critical role in diagnosis, treatment planning, and disease monitoring in neuro-oncology. Traditional methods of tumor segmentation, often manual and labour-intensive, are prone to inconsistencies and inter-observer variability. Recently, deep learning models, particularly Convolutional Neural Networks (CNNs), have shown great promise in automating this process. However, these models face challenges in terms of generalization across diverse datasets, accurate tumor boundary delineation, and uncertainty estimation. To address these challenges, we propose AG-MS3D-CNN, an attention-guided multiscale 3D convolutional neural network for brain tumor segmentation. Our model integrates local and global contextual information through multiscale feature extraction and leverages spatial attention mechanisms to enhance boundary delineation, particularly in complex tumor regions. We also introduce Monte Carlo dropout for uncertainty estimation, providing clinicians with confidence scores for each segmentation, which is crucial for informed decision-making. Furthermore, we adopt a multitask learning framework, which enables the simultaneous segmentation, classification, and volume estimation of tumors. To ensure robustness and generalizability across diverse MRI acquisition protocols and scanners, we integrate a domain adaptation module into the network. Extensive evaluations on the BraTS 2021 dataset and additional external datasets, such as OASIS, ADNI, and IXI, demonstrate the superior performance of AG-MS3D-CNN compared to existing state-of-the-art methods. Our model achieves high Dice scores and shows excellent robustness, making it a valuable tool for clinical decision support in neuro-oncology.

  • Research Article
  • Cite Count Icon 3
  • 10.3390/math12091281
MDER-Net: A Multi-Scale Detail-Enhanced Reverse Attention Network for Semantic Segmentation of Bladder Tumors in Cystoscopy Images
  • Apr 24, 2024
  • Mathematics
  • Chao Nie + 2 more

White light cystoscopy is the gold standard for the diagnosis of bladder cancer. Automatic and accurate tumor detection is essential to improve the surgical resection of bladder cancer and reduce tumor recurrence. At present, Transformer-based medical image segmentation algorithms face challenges in restoring fine-grained detail information and local boundary information of features and have limited adaptability to multi-scale features of lesions. To address these issues, we propose a new multi-scale detail-enhanced reverse attention network, MDER-Net, for accurate and robust bladder tumor segmentation. Firstly, we propose a new multi-scale efficient channel attention module (MECA) to process four different levels of features extracted by the PVT v2 encoder to adapt to the multi-scale changes in bladder tumors; secondly, we use the dense aggregation module (DA) to aggregate multi-scale advanced semantic feature information; then, the similarity aggregation module (SAM) is used to fuse multi-scale high-level and low-level features, complementing each other in position and detail information; finally, we propose a new detail-enhanced reverse attention module (DERA) to capture non-salient boundary features and gradually explore supplementing tumor boundary feature information and fine-grained detail information; in addition, we propose a new efficient channel space attention module (ECSA) that enhances local context and improves segmentation performance by suppressing redundant information in low-level features. Extensive experiments on the bladder tumor dataset BtAMU, established in this article, and five publicly available polyp datasets show that MDER-Net outperforms eight state-of-the-art (SOTA) methods in terms of effectiveness, robustness, and generalization ability.

  • Research Article
  • 10.1038/s41598-025-28315-9
Multi-scale aggregation network for colonoscopic polyp segmentation via frequency domain decoupling.
  • Dec 29, 2025
  • Scientific reports
  • Yanling Wang + 4 more

Automated segmentation of colorectal polyps is of great significance for early screening and clinical intervention of colorectal cancer. However, the diversity of polyp morphology and the uneven contrast caused by illumination changes in colonoscopy images make accurate segmentation and edge extraction of polyps a challenging task. To this end, this paper proposes a frequency domain decoupled multi-scale feature aggregation network (FDANet). The network employs wavelet transform to decompose spatial domain features into frequency domain sub-bands, extracting both low-frequency and high-frequency components. By leveraging their distinct frequency characteristics, the model is guided to suppress redundant information, emphasize target-relevant features, and achieve more robust and accurate segmentation results. In FDANet, the low-frequency attention enhancement module (LAEM) suppresses high-frequency background noise by performing Gaussian difference operations on low-frequency components and incorporates a hybrid attention mechanism to strengthen the feature representation of foreground regions. The high-frequency multi-scale aggregation module (HMAM) employs directional convolution kernels to model high-frequency components, extract fine-grained edge information, and construct a multi-scale feature pyramid to accommodate the morphological and scale diversity of polyps, while enhancing spatial detail awareness during the decoding stage. Additionally, an edge loss function is introduced to supervise the modeling of edge contours within this module, effectively suppressing background noise interference and further improving boundary localization accuracy. Experimental results show that this method achieves good segmentation results on CVC-ClinicDB and Kvasir-SEG datasets, outperforming other advanced segmentation methods.

  • Research Article
  • 10.3389/fphys.2025.1651296
MDWC-Net: a multi-scale dynamic-weighting context network for precise spinal X-ray segmentation
  • Aug 29, 2025
  • Frontiers in Physiology
  • Zhongzheng Gu + 2 more

PurposeSpinal X-ray image segmentation faces several challenges, such as complex anatomical structures, large variations in scale, and blurry or low-contrast boundaries between vertebrae and surrounding tissues. These factors make it difficult for traditional models to achieve accurate and robust segmentation. To address these issues, this study proposes MDWC-Net, a novel deep learning framework designed to improve the accuracy and efficiency of spinal structure identification in clinical settings.MethodsMDWC-Net adopts an encoder–decoder architecture and introduces three modules—MSCAW, DFCB, and BIEB—to address key challenges in spinal X-ray image segmentation. The network is trained and evaluated on the Spine Dataset, which contains 280 X-ray images provided by Henan Provincial People’s Hospital and is randomly divided into training, validation, and test sets with a 7:1:2 ratio. In addition, to evaluate the model’s generalizability, further validation was conducted on the Chest X-ray dataset for lung field segmentation and the ISIC2016 dataset for melanoma boundary delineation.ResultsMDWC-Net outperformed other mainstream models overall. On the Spine Dataset, it achieved a Dice score of 89.86% ± 0.356, MIoU of 90.53% ± 0.315, GPA of 96.82% ± 0.289, and Sensitivity of 96.77% ± 0.212. A series of ablation experiments further confirmed the effectiveness of the MSCAW, DFCB, and BIEB modules.ConclusionMDWC-Net delivers accurate and efficient segmentation of spinal structures, showing strong potential for integration into clinical workflows. Its high performance and generalizability suggest broad applicability to other medical image segmentation tasks.

  • Research Article
  • Cite Count Icon 30
  • 10.1109/jbhi.2021.3082527
Attention-RefNet: Interactive Attention Refinement Network for Infected Area Segmentation of COVID-19.
  • May 25, 2021
  • IEEE Journal of Biomedical and Health Informatics
  • Titinunt Kitrungrotsakul + 12 more

COVID-19 pneumonia is a disease that causes an existential health crisis in many people by directly affecting and damaging lung cells. The segmentation of infected areas from computed tomography (CT) images can be used to assist and provide useful information for COVID-19 diagnosis. Although several deep learning-based segmentation methods have been proposed for COVID-19 segmentation and have achieved state-of-the-art results, the segmentation accuracy is still not high enough (approximately 85%) due to the variations of COVID-19 infected areas (such as shape and size variations) and the similarities between COVID-19 and non-COVID-infected areas. To improve the segmentation accuracy of COVID-19 infected areas, we propose an interactive attention refinement network (Attention RefNet). The interactive attention refinement network can be connected with any segmentation network and trained with the segmentation network in an end-to-end fashion. We propose a skip connection attention module to improve the important features in both segmentation and refinement networks and a seed point module to enhance the important seeds (positions) for interactive refinement. The effectiveness of the proposed method was demonstrated on public datasets (COVID-19CTSeg and MICCAI) and our private multicenter dataset. The segmentation accuracy was improved to more than 90%. We also confirmed the generalizability of the proposed network on our multicenter dataset. The proposed method can still achieve high segmentation accuracy.

  • Research Article
  • Cite Count Icon 10
  • 10.1109/tmi.2024.3389776
Learning a Single Network for Robust Medical Image Segmentation With Noisy Labels.
  • Sep 1, 2024
  • IEEE transactions on medical imaging
  • Shuquan Ye + 4 more

Robust segmenting with noisy labels is an important problem in medical imaging due to the difficulty of acquiring high-quality annotations. Despite the enormous success of recent developments, these developments still require multiple networks to construct their frameworks and focus on limited application scenarios, which leads to inflexibility in practical applications. They also do not explicitly consider the coarse boundary label problem, which results in sub-optimal results. To overcome these challenges, we propose a novel Simultaneous Edge Alignment and Memory-Assisted Learning (SEAMAL) framework for noisy-label robust segmentation. It achieves single-network robust learning, which is applicable for both 2D and 3D segmentation, in both Set-HQ-knowable and Set-HQ-agnostic scenarios. Specifically, to achieve single-model noise robustness, we design a Memory-assisted Selection and Correction module (MSC) that utilizes predictive history consistency from the Prediction Memory Bank to distinguish between reliable and non-reliable labels pixel-wisely, and that updates the reliable ones at the superpixel level. To overcome the coarse boundary label problem, which is common in practice, and to better utilize shape-relevant information at the boundary, we propose an Edge Detection Branch (EDB) that explicitly learns the boundary via an edge detection layer with only slight additional computational cost, and we improve the sharpness and precision of the boundary with a thinning loss. Extensive experiments verify that SEAMAL outperforms previous works significantly.

  • Research Article
  • 10.1038/s41746-026-02371-5
Prompt-mamba filtering networks for accurate hepatocellular carcinoma lesion segmentation in abdominal CT.
  • Jan 27, 2026
  • NPJ digital medicine
  • Long Xia + 11 more

Precise delineation of hepatocellular carcinoma (HCC) in abdominal CT is pivotal for early diagnosis and surgical planning, yet remains challenged by morphological heterogeneity, low contrast in small lesions, and scanner variability. To address these limitations, we propose Prompt-Mamba-AF, a framework tailored for robust HCC segmentation. Our method uniquely integrates anatomy-aware prompts to guide feature extraction within liver regions and leverages Mamba-based state-space modeling to capture long-range volumetric dependencies with linear complexity. Furthermore, we introduce structure-aware filtering to enforce topological consistency along lesion boundaries. Extensive validation on the LiTS, 3DIRCADb, and CHAOS benchmarks demonstrates that Prompt-Mamba-AF outperforms current state-of-the-art CNN and Transformer architectures. The model achieves leading Dice similarity and boundary accuracy while maintaining a compact parameter footprint (27.6M). Results indicate significant improvements in small nodule sensitivity and generalization across diverse imaging domains, positioning Prompt-Mamba-AF as an efficient solution for multi-center clinical workflows.

  • Research Article
  • Cite Count Icon 24
  • 10.1016/j.eswa.2022.117360
DMDF-Net: Dual multiscale dilated fusion network for accurate segmentation of lesions related to COVID-19 in lung radiographic scans
  • May 2, 2022
  • Expert Systems with Applications
  • Muhammad Owais + 2 more

The recent disaster of COVID-19 has brought the whole world to the verge of devastation because of its highly transmissible nature. In this pandemic, radiographic imaging modalities, particularly, computed tomography (CT), have shown remarkable performance for the effective diagnosis of this virus. However, the diagnostic assessment of CT data is a human-dependent process that requires sufficient time by expert radiologists. Recent developments in artificial intelligence have substituted several personal diagnostic procedures with computer-aided diagnosis (CAD) methods that can make an effective diagnosis, even in real time. In response to COVID-19, various CAD methods have been developed in the literature, which can detect and localize infectious regions in chest CT images. However, most existing methods do not provide cross-data analysis, which is an essential measure for assessing the generality of a CAD method. A few studies have performed cross-data analysis in their methods. Nevertheless, these methods show limited results in real-world scenarios without addressing generality issues. Therefore, in this study, we attempt to address generality issues and propose a deep learning–based CAD solution for the diagnosis of COVID-19 lesions from chest CT images. We propose a dual multiscale dilated fusion network (DMDF-Net) for the robust segmentation of small lesions in a given CT image. The proposed network mainly utilizes the strength of multiscale deep features fusion inside the encoder and decoder modules in a mutually beneficial manner to achieve superior segmentation performance. Additional pre- and post-processing steps are introduced in the proposed method to address the generality issues and further improve the diagnostic performance. Mainly, the concept of post-region of interest (ROI) fusion is introduced in the post-processing step, which reduces the number of false-positives and provides a way to accurately quantify the infected area of lung. Consequently, the proposed framework outperforms various state-of-the-art methods by accomplishing superior infection segmentation results with an average Dice similarity coefficient of 75.7%, Intersection over Union of 67.22%, Average Precision of 69.92%, Sensitivity of 72.78%, Specificity of 99.79%, Enhance-Alignment Measure of 91.11%, and Mean Absolute Error of 0.026.

  • Research Article
  • Cite Count Icon 20
  • 10.1007/s10489-021-02297-3
Adaptive channel and multiscale spatial context network for breast mass segmentation in full-field mammograms
  • Apr 14, 2021
  • Applied Intelligence
  • Wenwei Zhao + 6 more

Breast cancer is currently the second most fatal cancer in women, but timely diagnosis and treatment can reduce its mortality. Breast masses are the most obvious means of cancer identification, and thus, accurate segmentation of masses is critical. In contrast to mass-centered patch segmentation, accurate segmentation of breast masses in full-field mammograms is always a challenging topic because of the extremely low signal-to-noise ratio and the uncertainty with respect to the shape, size, and location of the mass. In this study, we propose a novel adaptive channel and multiscale spatial context network for breast mass segmentation in full-field mammograms. A standard encoder-decoder structure is employed, and an elaborate adaptive channel and multiscale spatial context module (ACMSC module) is embedded in a multilevel manner in our network for accurate mass segmentation. The proposed ACMSC module utilizes the self-attention mechanism to adaptively capture discriminative contextual information among channel and spatial dimensions.The multilevel embedding of the ACMSC module enables the network to learn distinguishing features on multiple scales of feature maps. Our proposed model is evaluated on two public datasets, CBIS-DDSM and INbreast. The experimental results show that by adaptively capturing the context of the channel and spatial dimensions, our model can effectively remove false positives, predict difficult samples and achieve state-of-the-art results, with Dice coefficients of 82.81% for CBIS-DDSM and 84.11% for INbreast, respectively. We hope that our work will contribute to the CAD system for breast cancer diagnosis and ultimately improve clinical diagnosis.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant