Articles published on Semantic enhancement
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
359 Search results
Sort by Recency
- New
- Research Article
- 10.1080/17538947.2026.2625542
- Jul 1, 2026
- International Journal of Digital Earth
- Xiao Ling + 6 more
Accurate information on rice planting areas and spatial distribution is critical for agricultural management in China; however, mapping efforts in regions like Sichuan Province are severely constrained by persistent cloud cover and fragmented terrain. Existing phenology-based methods and coarse-resolution products often fail to provide precise paddy localizations or usable training labels in such complex environments. To address these limitations, this study proposes a robust framework integrating optical and Synthetic Aperture Radar (SAR) imagery. The methodology employs a phenology-driven strategy for rapid candidate area annotation, coupled with an asymmetric feature extraction mechanism that incorporates a Multi-Scale Semantic Enhancement Module (MSEM) and a Category-Balanced Feature Fusion Module (CBFM) to facilitate adaptive and effective cross-modal fusion. Applying this framework to the Tianfu New Area (2019–2023) yielded 10-meter resolution rice distribution maps with a rice Intersection over Union (IoU) of 83.31% and a statistical correlation ( R 2 ) of 0.972. These results demonstrate the framework’s capacity for selective multi-source fusion and cost-effective sampling, facilitating precise rice mapping in challenging agricultural landscapes.
- Research Article
- 10.3390/s26102956
- May 8, 2026
- Sensors (Basel, Switzerland)
- Boning Zhu + 2 more
Vision foundation models (VFMs) enable high-quality 2D instance masks, yet their lifted pseudo-point clouds suffer from scale ambiguity, structural noise, and temporal inconsistency, limiting their utility in 3D annotation. Existing automatic labeling methods either rely on expensive light detection and ranging (LiDAR) sensors or fail to enforce physical plausibility in dynamic roadside scenes. This study proposes a LiDAR-free radar–visual auto-labeling framework that leverages cross-modal spatio-temporal consistency between millimeter-wave radar trajectories and visual pseudo-point clouds to self-correct 3D geometry. The method first associates radar points, 2D masks, and pseudo-point clouds into object-centric sequences. Then, an uncertainty-aware pose fusion module combines motion-derived and structure-derived orientations using automatically solved road priors. Finally, the pseudo-point cloud is refined in canonical space by optimizing stable semantic landmarks from temporally consistent masks and propagating their corrections globally. Evaluated on a real-world roadside dataset, the method achieves 49.1% bird’s-eye-view (BEV) intersection over union (IoU) and 43.0% 3D IoU, outperforming a radar–camera fusion baseline by 5.5/5.9 points. Downstream experiments further show that the generated pseudo-labels and semantic enhancement are useful under the evaluated detector configurations, while broader validation remains future work.
- Research Article
- 10.3390/s26092864
- May 3, 2026
- Sensors (Basel, Switzerland)
- Yawen Zhu + 6 more
Withthe widespread deployment of intelligent terminals, mobile payment platforms, and Internet of Things devices, security systems are being progressively transformed from traditional transaction outcome analysis toward an intelligent perception paradigm centered on user behavior, device states, and environmental context. To address the challenges of multimodal data heterogeneity, non-independent and identically distributed data across nodes, and the difficulty of centralized modeling under privacy constraints in distributed scenarios, an artificial intelligence-driven federated multimodal security perception framework, namely FMS-LLM, is proposed. At its core, the framework introduces a Non-IID adaptive federated fusion mechanism that achieves dual-level alignment—structural alignment via parameter-level masks and semantic alignment via feature consistency constraints—to effectively mitigate cross-node distribution discrepancies. Additionally, an LLM-driven semantic enhancement module is developed, utilizing trend-guided token selection and inertia-suppression to map low-level sensing features into high-level risk semantic representations, thereby supporting logical reasoning and explainable decision-making. This framework takes user behavioral sensing data, device state information, environmental context data, and transaction behavior data as inputs, and constructs an integrated security analysis pipeline of “perception–collaboration–reasoning”. Experimental results on the distributed multimodal security perception task demonstrate that the proposed method achieves an Accuracy of 91.62%, a Precision of 91.04%, a Recall of 90.37%, an F1-score of 90.70%, and a ROC-AUC of 94.73%, consistently outperforming baseline methods including Logistic Regression, Random Forest, LSTM, the centralized multimodal deep model, FedAvg, FedProx, and MOON. Under strongly Non-IID conditions, when , the model still maintains an Accuracy of 88.47% and an F1-score of 87.11%, demonstrating stronger cross-node robustness. The ablation study further indicates that the complete model attains the best classification performance while reducing communication cost to 18.92 MB/Round. These results demonstrate that the proposed method can effectively fuse multi-source sensing information under privacy-preserving conditions and support intelligent security perception tasks with higher accuracy, stronger robustness, and improved interpretability.
- Research Article
- 10.1016/j.eswa.2026.131390
- May 1, 2026
- Expert Systems with Applications
- Zhijing Yang + 7 more
Neural clothing tryer: Customized virtual try-on via semantic enhancement and controlling diffusion model
- Research Article
- 10.1038/s41598-026-46407-y
- Apr 25, 2026
- Scientific reports
- Jia Yu + 8 more
Hyperspectral image classification has become a pivotal task in remote sensing data processing. However, current deep learning-based methods still face challenges in effectively extracting discriminative joint spatial-spectral features, especially when handling complex spatial structures and highly variable spectral responses. To address this issue, a novel hyperspectral image classification network (MSNet) based on multiscale spatial-spectral fusion (MSSF) and semantic enhancement encoder (SEE) is proposed. First, principal component analysis is employed to reduce the dimensionality of hyperspectral image, preserving essential spectral information while mitigating noise interference. Second, the proposed MSSF extracts multiscale spatial features through a dedicated spatial branch while simultaneously mining discriminative spectral information via a parallel spectral branch. Feature fusion is subsequently employed to achieve synergistic interaction between spatial and spectral representations, effectively capturing joint characteristics of land covers across multiple scales. Third, the designed SEE incorporates a multi-head attention mechanism to explicitly model global feature dependencies, thereby enhancing the representation of semantically critical regions. Extensive experiments have been conducted on the Pavia University and Salinas datasets. Experimental results demonstrate that MSNet achieves outstanding performance, with overall accuracy reaching 95.68% and 96.84%, respectively, surpassing existing mainstream methods and validating its effectiveness.
- Research Article
- 10.1038/s41598-026-49439-6
- Apr 24, 2026
- Scientific reports
- Lei Wang + 7 more
PFL-YOLO: a robust framework for lightweight product damage detection with multi-scale semantic enhancement in unmanned vending systems.
- Research Article
- 10.3390/a19050325
- Apr 22, 2026
- Algorithms
- Yujie Du + 3 more
Semantic segmentation of land use and land cover (LULC) in arid regions remains challenging due to severe class imbalance, fragmented spatial distributions, and high spectral similarity among different land cover types. These characteristics often lead to an information bottleneck in deep segmentation networks and hinder the extraction of discriminative semantic representations. To address these issues, we propose SDS-Former, a lightweight semantic segmentation network specifically designed for remote sensing imagery in arid environments. SDS-Former incorporates an SSM-inspired Lightweight Semantic Enhancement (LSE) module to strengthen contextual modeling and alleviate the loss of discriminative information in deep features. To tackle scale variations, a Dynamic Selective Feature Fusion (DSFF) module is employed in the decoder to adaptively weight and fuse high-level semantics with low-level spatial details. Furthermore, a Feature Refinement Head (FRH) is introduced to enhance boundary localization and improve the recognition of small-scale and sparsely distributed land cover objects. Extensive ablation and comparative experiments demonstrate that SDS-Former consistently outperforms representative semantic segmentation methods across multiple evaluation metrics. On the Tarim Basin dataset, the proposed network achieves a mean Intersection over Union (mIoU) of 82.51% and an F1 score of 86.47%, indicating its superior effectiveness and robustness. Qualitative results further verify that SDS-Former exhibits clear advantages in distinguishing spectrally similar land cover types and preserving the spatial continuity of ground objects in complex arid-region scenes.
- Research Article
- 10.1080/13682199.2026.2658400
- Apr 17, 2026
- The Imaging Science Journal
- Ganesh Khekare + 4 more
ABSTRACT Low visibility, color distortion, and structural complexity are some of the harsh challenges that marine environments must confront. This research affords a transformational image caption structure, specifically designed for the assessment of underwater sceneries. To produce precise and significant captions for underwater images, the suggested technique combines seen-linguistic fusion, contextual and semantic enhancements, and hobby mechanisms. This proposed method consists of unique local language and spatial skills, enabling greater unique interpretation and context identity of complicated scenarios. The model achieves an increased accuracy of 91.40%, which extensively outperforms present techniques. The proposed technique consists of cutting-edge measures like Bilingual Evaluation Understudy (BLEU), Metric for Evaluation of Translation with Explicit ORdering (METEOR) and Consensus-based Image Description Evaluation (CIDEr). Contrast analysis and case studies show that the system may create captioning that is both aesthetically pleasing and linguistically rich, making it a valuable tool for tracking, exploration, and marine recording.
- Research Article
- 10.1038/s41598-026-46737-x
- Apr 7, 2026
- Scientific reports
- Jing Dong + 4 more
With the increasing challenge of information overload, personalized recommendation systems play an essential role in delivering relevant content to users. Traditional collaborative filtering methods often suffer from data sparsity and cold-start problems, which limit their effectiveness in real-world applications. To address these issues, this paper proposes a personalized recommendation model that integrates graph attention mechanisms with auxiliary textual information. User-item interactions are modeled on a knowledge graph, where graph attention networks are employed to capture multi-hop user interests, while graph convolution is used to aggregate item-side neighborhood information. Textual data associated with users and items are encoded into semantic embeddings and incorporated to enrich the initial representations of graph entities. Experiments conducted on the Book-Crossing and MovieLens-1M datasets demonstrate that the proposed model achieves superior performance in terms of AUC, F1-score, and Top-K recall compared with several state-of-the-art baselines. The results indicate that combining graph-based modeling with textual semantic enhancement can effectively improve recommendation accuracy and robustness under sparse data conditions.
- Research Article
3
- 10.1109/tfuzz.2025.3559562
- Apr 1, 2026
- IEEE Transactions on Fuzzy Systems
- Zhan Gao + 5 more
Early disease diagnosis is critical for timely clinical intervention and treatment. Intelligent models have shown significant potential in addressing the challenges of misdiagnosis, especially given the shortage of experienced experts. However, there exists complexity of ultrasound image information, and subtle differences between positive and negative samples, combined with limited disease data, pose challenges in feature extraction, class imbalance, and the lack of representative positive class prototypes. To this end, we propose a large vision model distillation framework with fuzzy perception to improve rare disease diagnosis. Specifically, we first fine-tune the pre-trained Medical Segment Anything Model (MedSAM) to adapt it to the target domain. Through image augmentation and latent feature relationship distillation, we enhance feature extraction robustness and reduce inter-class ambiguity, which helps mitigate the impact of class imbalance on the lightweight student model. Second, as imaging style differences in ultrasound images are often more pronounced than subtle variations between positive and negative samples, we construct image style subsets and introduce a fuzzy style matching strategy to perceive these differences. Finally, we combine features from both the differential perception and semantic enhancement branches to strengthen disease classification. Extensive experiments on an internal fetal spina bifida dataset and three widely used imbalanced benign-malignant medical datasets demonstrate the effectiveness of the proposed method.
- Research Article
- 10.1016/j.neunet.2025.108357
- Apr 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Yawen Shen + 2 more
Improving the robustness of graph contrastive learning against adversarial attacks via hierarchical medoid-based contrasting.
- Research Article
- 10.3390/electronics15071446
- Mar 30, 2026
- Electronics
- Ruoxi Liu + 2 more
Temporal Knowledge Graph Inference (TKGI) is a cornerstone for intelligent decision-making in dynamic scenarios, but existing models face critical bottlenecks, including inadequate complex-context modeling, a lack of entity importance quantification, insufficient novel-event reasoning accuracy, and weak domain adaptability. To address these issues, this study proposes a semantics-enhanced model (LLM-DSaR) integrating Large Language Models (LLMs), temporal attention networks, and optimized contrastive learning. Specifically, a two-stage LLM semantic enhancement (LLM1 + LLM2) framework first generates structured semantic analysis reports via adaptive prompt engineering, and then extracts domain-specific semantic embeddings from the last-layer hidden states through pooling and linear projection, which are further fused with TransE-based structural embeddings; meanwhile, LLM2 mitigates data sparsity in novel-event reasoning; a dynamic weight fusion (DWF) framework adaptively assigns feature weights to achieve deep feature synergy; an LLM-enhanced contrastive-learning module strengthens event clustering and discrimination. Experiments on five public datasets and a self-constructed Robotics Temporal Knowledge Graph (RTKG) show LLM-DSaR outperforms 16 baselines: on RTKG, its MRR is 10.35 percentage points higher than GCR, and Hits@10 reaches 88.87%. Ablation experiments validate core modules’ effectiveness, confirming LLM-DSaR adapts to professional scenarios like robot maintenance prediction, providing a novel technical paradigm for complex-domain TKG reasoning.
- Research Article
- 10.3390/app16062943
- Mar 18, 2026
- Applied Sciences
- Tianzhe Jiao + 4 more
Multimodal fusion methods leveraging various sensors provide strong support for 3D object detection. However, under adverse weather conditions such as rain, fog, snow, and intense glare, complex environmental factors can degrade sensor data quality, leading to increased false positives and missed detections. In addition, sensor modalities (e.g., LiDAR and cameras) inherently vary in information density, and directly fusing them can cause critical details in high-density data to be diluted by low-density data, thereby increasing errors. To address these issues, we propose a Semantic-Enhanced Bidirectional Multimodal Fusion (SeBFusion) framework. By introducing a semantic enhancement mechanism and a bidirectional fusion strategy, SeBFusion mitigates the impact of noise under adverse weather and alleviates information dilution in multimodal fusion. Specifically, SeBFusion first employs a virtual point generation and camera semantic injection module to selectively map image semantic features into 3D space, producing semantically enhanced LiDAR features to compensate for the sparsity of the raw LiDAR point cloud. Then, during cross-modal interaction, we design a bidirectional cross-attention fusion module. This module estimates the confidence of each modality and adaptively reweights the bidirectional information flow, thereby reducing the risk of noise propagation across modalities and improving the robustness and accuracy of 3D object detection in complex environments. Experiments on adverse-weather versions of datasets such as KITTI-C and nuScenes-C validate the effectiveness and superiority of the proposed method. On the nuScenes-C dataset, it achieves 66.2% mAP and 66.6% mAP under fog and snow conditions, respectively.
- Research Article
- 10.54097/825wrs61
- Mar 8, 2026
- International Journal of Advanced Engineering and Technology Research
- Daini Li + 1 more
The incidence of auricular deformities in newborns is notably high, and even experienced clinicians may encounter issues such as misdiagnosis and missed diagnosis due to subjective judgment. Although several studies have explored the use of deep learning methods for auxiliary diagnosis, the highly complex and individualized characteristics of auricular morphology pose significant challenges to existing approaches in achieving automated identification and fine-grained subtype classification. To address this issue, we propose HRS-Swin, a progressive representation reconstruction framework built upon a Swin Transformer backbone. The model integrates a Class Token Fusion module to enhance global semantic representation, a Stable Semantic Enhancement and Residual Compression mechanism for compact and discriminative embedding learning, and a Dynamic Margin Enhancer to enlarge inter-class separability in the embedding space. Experiments on the BabyEar4K dataset (1,926 newborns) demonstrate that HRS-Swin outperforms representative CNN and Transformer baselines. The proposed method achieves an accuracy of 0.8009 and a macro F1-score of 0.7024, showing consistent improvements over standard Swin Transformer. These results indicate that the proposed framework provides a robust and effective solution for automated auricular deformity classification and early clinical assistance.
- Research Article
1
- 10.1016/j.eswa.2025.129857
- Mar 1, 2026
- Expert Systems with Applications
- Zhaoyi Meng + 4 more
SmartScope: Smart contract vulnerability detection via heterogeneous graph embedding with local semantic enhancement
- Research Article
- 10.1016/j.ins.2025.122882
- Mar 1, 2026
- Information Sciences
- Jinlin Jiang + 3 more
DSWFusion: Separation-guided multi-frequency semantic enhancement model for infrared and visible image fusion
- Research Article
- 10.1016/j.isprsjprs.2026.01.017
- Mar 1, 2026
- ISPRS Journal of Photogrammetry and Remote Sensing
- Kai Hu + 5 more
Knowledge distillation with spatial semantic enhancement for remote sensing object detection
- Research Article
2
- 10.62762/tis.2025.879161
- Feb 19, 2026
- ICCK Transactions on Intelligent Systematics
- Asad Ullah Haider + 2 more
Defocus blur detection is essential for computational photography applications, but existing methods struggle with accurate blur localization and boundary preservation. We propose SemanticBlur, a deep learning framework which integrates semantic understanding with attention mechanisms for robust defocus blur detection. Our semantic-aware attention module combines channel attention, spatial attention, and semantic enhancement to leverage high-level features for low-level feature refinement. The architecture employs a modified ResNet-50 backbone with dilated convolutions that preserves spatial resolution while expanding receptive fields, coupled with a feature pyramid decoder using learnable fusion weights for adaptive multi-scale integration. A combined loss function balancing binary cross-entropy and structural similarity achieves both pixel-wise accuracy and structural coherence. Extensive experiments on four benchmark datasets (CUHK, DUT, CTCUG, EBD) demonstrate state-of-the-art (SOTA) performance, with ablation studies confirming that semantic enhancement provides the most significant gains while maintaining computational efficiency. SemanticBlur generates visually coherent detection maps with sharp boundaries, validating its practical applicability for real-world deployment.
- Research Article
- 10.70849/ijsci03022639353
- Feb 13, 2026
- International Journal of Sciences and Innovation Engineering
- Sakshi Gupta
This research offers an in-depth semantic enhancement and structural analysis of three key cloud ontologies, CloudLightning, CoCoOn, and Mosaic. Employing RDFLib and Python-driven analytics, we methodically assessed and enhanced the ontologies via annotation enrichment, namespace standardization, and structural verification. The enhanced ontologies show notable advancements in semantic precision, interoperability, and annotation scope. Quantitative findings indicate rises of more than 200 new classes and over 300 annotation enhancements throughout the three ontologies. The results emphasize that structured improvement boosts ontology development, allowing for improved reasoning and interoperability among cloud systems.
- Research Article
- 10.1038/s41598-026-38628-y
- Feb 5, 2026
- Scientific reports
- Lujin Zhao + 6 more
Cross-domain sequential recommendation (CDSR) models users’ dynamic preferences by exploiting behavioral signals from multiple domains, but it faces challenges in data sparsity, domain heterogeneity, and privacy protection. Although federated learning enables privacy-preserving CDSR by keeping raw data local, existing methods often suffer from sparse representations, unstable cross-domain alignment, and severe utility degradation under uniform differential privacy. In this work, we propose FedSCOPE, a novel federated CDSR framework that addresses these challenges through three tightly coupled and explicitly aligned components. First, FedSCOPE enriches user and item representations via offline large language model (LLM)-generated semantic augmentation, mitigating sparsity while avoiding online LLM inference and the associated privacy and deployment risks. Second, it introduces an Intra- and Inter-Domain Decoupled Contrastive Learning mechanism that separates intra-domain personalization from inter-domain discrimination, enabling robust cross-domain alignment under heterogeneous data distributions. Third, FedSCOPE incorporates an adaptive personalized differential privacy strategy that dynamically allocates privacy budgets and clipping thresholds according to client-specific data characteristics, achieving a more favorable privacy–utility trade-off in federated environments. These components are jointly optimized within a secure federated learning framework. Extensive experiments on multiple real-world datasets demonstrate that FedSCOPE consistently outperforms state-of-the-art baselines, achieving higher recommendation accuracy, stronger cross-domain generalization, and improved privacy–utility balance.