Articles published on Noisy data
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
12177 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.cscm.2025.e05671
- Jul 1, 2026
- Case Studies in Construction Materials
- Maryam Abazarsa + 1 more
A deep learning approach for predicting steel rebar corrosion in concrete bridge columns from two-year noisy GPR B-scan images
- New
- Research Article
- 10.1016/j.neucom.2026.133614
- Jul 1, 2026
- Neurocomputing
- Chengming Liu + 3 more
A synergy scoring filter for unsupervised anomaly detection with noisy data
- New
- Research Article
- 10.1021/acs.jctc.6c00372
- Jul 1, 2026
- Journal of chemical theory and computation
- Jo S Kurian + 2 more
In this article, we present a method for computing accurate and scalable nuclear forces within the phaseless auxiliary-field quantum Monte Carlo (AFQMC) framework. Our approach leverages automatic differentiation of the energy functional to obtain nuclear gradients at a computational cost comparable to that of energy evaluation. The accuracy of the method is validated against finite difference calculations, showing excellent agreement. We then explore several machine learning (ML) strategies for learning noisy AFQMC data. These ML potentials are subsequently used to perform geometry optimizations and nudged elastic band (NEB) calculations, successfully identifying the transition state of the formamide-formimidic acid tautomerization. The resulting transition-state geometry and barrier heights are in close agreement with coupled-cluster reference values. This work paves the way for highly accurate geometry optimization, molecular dynamics, or reaction path calculations.
- New
- Research Article
- 10.1002/mrm.70312
- Jul 1, 2026
- Magnetic resonance in medicine
- Daiki Tamada + 4 more
To introduce and evaluate the feasibility of a novel RF-phase modulated gradient echo (GRE) method for quantitative diffusion MRI, aimed at mitigating geometric distortion and enabling high-resolution 3D quantitative diffusion/T2 mapping as a complementary alternative to conventional DWI. The proposed phase-based diffusion (PBD) method employs RF phase modulation to encode both diffusion and T2 information into the GRE signal phase. A closed-form analytical model enables joint apparent diffusion coefficient (ADC) and T2 mapping via iterative reconstruction. The method's feasibility was evaluated via Bloch equation simulations, phantom experiments, and preliminary in vivo imaging studies. Monte Carlo simulations revealed that PBD provides more accurate median ADC estimates at low signal-to-noise ratios (SNRs) compared to conventional single-shot echo-planar imaging (SS-EPI), although PBD exhibited greater variability. Phantom studies demonstrated good agreement for PBD-derived ADC values (e.g., R2 = 0.99) with reference methods and strong correlation for PBD-derived T2 values (e.g., R2 = 0.89), though the latter showed some systematic bias in phantoms. In vivo results from patients with benign or malignant prostate disease demonstrated the feasibility of the PBD method to provide high-resolution ADC and T2 maps with minimal geometric distortions relative to conventional SS-EPI. PBD provides ADC and T2 maps with improved geometric fidelity in phantoms and in vivo, and offers robust median ADC estimates from noisy data based on simulations. This combination of spatial precision and noise characteristics makes PBD promising for applications such as high-resolution DWI for prostate MRI.
- New
- Research Article
- 10.1038/s41598-026-59117-2
- Jun 29, 2026
- Scientific reports
- Alawi A Al-Saggaf + 1 more
Many practical security applications rely on inherently noisy data sources, such as biometric measurements, sensor-derived features, and physical unclonable functions (PUFs). This variability makes it difficult to develop reliable cryptographic primitives. Conventional solutions-like fuzzy commitment schemes, fuzzy vaults, secure sketches, and fuzzy extractors-enable key binding or generation from noisy data by combining cryptography with error-correcting codes to handle variations. However, using ECCs can cause information leakage. In this paper, we construct ISNR-PQC, an isometry noise resilient post-quantum cryptographic primitive designed for natural noisy sources, using secure biometric data as an application. ISNR-PQC specifically uses lattice-based cryptography based on the National Institute of Standards and Technology (NIST) standard FIPS 203. Our approach provides security guarantees against quantum adversaries under the Learning with Errors (LWE) assumption. We formalise correctness, robustness, and [Formula: see text]-restricted IND-CPA indistinguishability in a unified security model. Our ISNR-PQC is noise resilient because it accepts ancillary data that is close to the original encrypting data, as measured by an isometry-based threshold mapping. Our results demonstrate that ISNR-PQC addresses a major challenge in biometric authentication theory and is a promising base for future biometric and noisy-data security systems.
- New
- Research Article
- 10.1088/1361-6560/ae792b
- Jun 25, 2026
- Physics in Medicine & Biology
- Haoyang Pei + 8 more
Objective.Deep learning has demonstrated strong potential for magnetic resonance imaging (MRI) reconstruction. However, conventional supervised learning requires high-quality, high-signal-to-noise-ratio (SNR) reference data for network training, which are often difficult or impossible to obtain, particularly in low-field MRI. Self-supervised learning (SSL) eliminates the need for reference training data but may suffer from degraded performance under low-SNR conditions. To address these limitations, we propose hybrid learning, a new training framework that integrates self-supervised and supervised learning for joint MRI reconstruction and denoising when only low-SNR training data are available.Approach.Hybrid learning is implemented in two sequential stages. In the first stage, SSL is applied to fully sampled low-SNR data to generate higher-quality pseudo-references. In the second stage, these pseudo-references are then used as targets for supervised learning to reconstruct and denoise undersampled, noisy data. The proposed method was evaluated in four experiments using simulated and real noisy MRI data of the breast, lung, and brain across different field strengths (0.3 T to 3 T), sampling trajectories (Cartesian, spiral, and radial), noise levels, and undersampling ratios.Main Results.Hybrid learning consistently improved reconstruction quality relative to both supervised and self-supervised baselines under different acceleration rates, noise levels, and sampling patterns in all experiments. Compared with standard supervised learning using noisy references, it achieved up to 167.70% higher structural similarity index measure (SSIM), 95.41% lower normalized mean squared error (NMSE), and 90.70% lower high-frequency error norm (HFEN). Compared with standard SSL, it achieved up to 23.88% higher SSIM, 60.85% lower NMSE, and 49.13% lower HFEN.Significance.Hybrid learning enables improved MRI reconstruction under low-SNR imaging conditions by jointly addressing noise and undersampling. It provides a practical solution for robust deep learning-based reconstruction and is particularly well suited for applications such as low-field MRI, where image quality is limited by reduced SNR.
- New
- Research Article
- 10.1016/j.ultras.2026.108206
- Jun 24, 2026
- Ultrasonics
- Hongsheng Xu + 4 more
The convergence of surface acoustic wave technology and artificial intelligence.
- New
- Research Article
- 10.1371/journal.pone.0352016
- Jun 23, 2026
- PLOS One
- Lei Ren + 1 more
This paper introduces a novel approach using physics-informed neural networks (PINNs) to simultaneously solve variable-order time fractional diffusion equations and infer the time-dependent fractional order from data. By embedding the governing equations into the neural network’s loss function, our method achieves high accuracy and flexibility, even with sparse or noisy data. We present a dual-network architecture where one network approximates the solution u(x,t) while another learns the fractional order . Numerical experiments demonstrate the effectiveness of our approach, achieving mean squared errors below 10−4 for solutions and 10−3 for fractional orders in smooth cases, while also handling noisy data and non-smooth orders robustly.
- New
- Research Article
- 10.1158/0008-5472.can-25-4371
- Jun 22, 2026
- Cancer research
- David E Frankhouser + 18 more
Single-cell RNA-sequencing (scRNA-seq) has revolutionized our understanding of cancer. However, identifying meaningful disease states from single-cell data remains challenging due to the complex continuum of transcriptional alterations, which may obscure clear phenotype boundaries and complicate biological interpretation and clinical relevance. Here, we systematically explored the chronic myeloid leukemia (CML) specific information content encoded in scRNA-seq versus bulk transcriptomics to resolve this paradox and clarify how discrete disease-defining states emerge from inherently noisy single-cell data. While CML single-cell transcriptomes existed along continuous transcriptional micro-states, clinically relevant leukemia phenotypes clearly manifested only at the pseudobulk (macro-state) level. State-transition theory was leveraged to reveal how disease phenotype state-transitions are governed by cell type specific contributions. Together, these results establish a theoretical framework explaining why discrete disease phenotypes remain hidden at the single-cell scale but emerge clearly at the aggregated macro-state level, enabling previously inaccessible biological insights into leukemia evolution. By resolving how single-cell variation aggregates into macroscopic disease states, this framework provides insights into CML progression and offers a broadly applicable strategy for exploring disease dynamics across cancers and other complex conditions.
- New
- Research Article
- 10.1038/s41598-026-58276-6
- Jun 20, 2026
- Scientific reports
- Akbar Hussain + 7 more
Accurate prediction of diseases from multiple co-occurring symptoms remains an important challenge in intelligent healthcare systems. Machine-learning models can support early symptom-based risk screening; however, their clinical use is limited by dataset quality, lack of external validation, interpretability concerns, and the risk of overstated diagnostic claims. This study presents a mobile decision-support prototype for multi-symptom disease prediction and rule-based lifestyle recommendation. The experiments were conducted using the publicly available Kaggle Medicine Recommendation System Dataset, which contains 4,920 symptom-based records spanning 41 disease classes, along with supporting files for disease descriptions, precautions, medications, dietary suggestions, and workout recommendations. Seven supervised machine-learning models were evaluated, including Decision Tree, Random Forest, Naive Bayes, Logistic Regression, XGBoost, Support Vector Machine, and K-Nearest Neighbors. The best-performing models achieved high classification accuracy under the adopted evaluation protocol. To avoid overinterpretation, these results are reported as benchmark performance on a secondary public dataset rather than evidence of clinical diagnostic validity. The Android-Flask prototype links the predicted disease class to a lookup-based recommendation layer that retrieves disease-associated information from supporting datasets. The system should therefore be interpreted as a decision-support and educational prototype, not as a clinically validated diagnostic or prescribing tool. Future work should include external validation in independent clinical cohorts, clinician assessment of recommendations, robustness testing under missing or noisy symptom data, and broader evaluation across real-world healthcare settings.
- New
- Research Article
- 10.1037/met0000842
- Jun 18, 2026
- Psychological methods
- Mischa Von Krause + 1 more
Recent advances in Bayesian modeling and deep learning have enabled scalable estimation of cognitive process models. In this article, we present a fully Bayesian workflow that leverages amortized inference with neural networks to rapidly estimate individual parameters and compare models from big behavioral data. Using data from a large online implicit association test sample (N > 5,000,000), we investigate how latent parameters, such as drift rate, boundary separation, nondecision times, and their variabilities, relate to key socioeconomic variables. Our exploratory findings reveal small but consistent associations of cognitive model parameters with socioeconomic covariates. Notably, trial-by-trial variability in drift rate, often ignored in prior work, emerged as the strongest predictor across all socioeconomic covariates. Our primary contribution lies in illustrating how deep learning-based Bayesian estimation and model comparison can be applied to mine robust insights from large and noisy behavioral data sets. We discuss limitations and implications for modeling individual differences in large-scale data sets and provide an open pipeline for future use. This work exemplifies how the emerging field of behavioral data science can extend cognitive modeling to new domains and support data-driven hypothesis generation targeting the cognitive underpinnings of individual differences. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
- New
- Research Article
- 10.1209/0295-5075/ae723c
- Jun 16, 2026
- Europhysics Letters
- Xinyi Gu + 2 more
Convergent Cross-Mapping (CCM) is a widely used nonlinear approach for detecting causality in complex systems, with applications in ecology, neuroscience, and economics. Traditionally, CCM employs Euclidean distance for nearest-neighbor identification, but this choice can be limited in high-dimensional, noisy, or scale-heterogeneous data. To address this, we systematically compare six metrics —Euclidean, Standardized Euclidean, Manhattan, Chebyshev, Cosine, and Correlation— within the CCM framework. Using coupled Logistic and Lorenz-Rössler systems, we assess their ability to capture causal dynamics under weak, strong, and asymmetric coupling. We further apply them to empirical data of neuronal multi-unit activity and behavioral rhythms in rodents, highlighting their performance in real-world conditions. Results reveal that no single distance metric is universally optimal: Chebyshev distance performs well for short sequences; Cosine distance excels in strong symmetric coupling; Standardized Euclidean is robust to scale differences and noise; Euclidean remains stable across settings; while Correlation distance underperforms due to its linearity. These findings underscore the importance of distance selection in CCM and provide methodological guidance for diverse applications.
- Research Article
- 10.1016/j.ultras.2026.108188
- Jun 11, 2026
- Ultrasonics
- Boyoung Kim + 1 more
Computational framework for guided wave NDT to identify delamination of arbitrary geometries in 3D layered anisotropic composite structures via machine learning.
- Research Article
- 10.1038/s41598-026-56254-6
- Jun 11, 2026
- Scientific reports
- Ganesh Karthikeyan Varadarajan + 6 more
The rapid advancement of Artificial Intelligence (AI) in healthcare diagnostics has significantly improved disease detection and treatment prediction; however, misclassification remains a persistent challenge due to data heterogeneity, incomplete clinical information, and complex feature interdependencies. To address these limitations, a novel framework titled Dual Motif Guided Heterogeneous Interactive Dandelion Graph Multi-Scale Residual Attention Network (DMGH-IDGM-SRAN) is proposed for Reducing Misclassifications and Enhancing Model Performance in Healthcare AI. The study employs a curated dataset, "Healthcare AI Dataset for Reducing Misclassifications and Enhancing Model Performance in Healthcare AI", collaboratively developed by technical and medical experts, consisting of 743 clinical records obtained from the Pulmonology Department of one of India's leading hospitals. Pre-processing is conducted using the Fuzzy Min-Max Neural Network (FMNN) to handle uncertainty and overlapping class boundaries effectively. Subsequently, Adaptive Causal Decision Transformers (AdaCred) are utilized for feature extraction, capturing causal dependencies and temporal associations across heterogeneous clinical attributes. The proposed DMGH-IDGM-SRAN integrates a Dual-Attention-Guided Interactive Multi-Scale Residual Network (DA-IMRN) with a Motif-Based Heterogeneous Graph Attention Network (MBHAN), whose hyperparameters are fine-tuned using the Dandelion Optimizer (DO) to enhance model convergence and stability. Experimental evaluation demonstrates that the proposed model achieves an exceptional 99.9% classification accuracy, significantly outperforming existing approaches. The framework offers two key advantages: (i) it effectively captures complex cross-feature relationships between clinical and biochemical parameters, and (ii) it provides robust and interpretable diagnostic predictions under noisy and incomplete data conditions.
- Research Article
- 10.1039/d6sm00198j
- Jun 10, 2026
- Soft matter
- Babak Valipour Goodarzi + 1 more
Elastoviscoplastic (EVP) materials pose significant challenges for predictive modeling, especially under large amplitude oscillatory shear (LAOS) conditions due to their nonlinear rheological behavior. While various models have been proposed to describe yielding and viscoelastic transitions, their application to experimental data remains limited due to computational complexity, poor generalizability, and difficulty in fitting noisy data. This work re-evaluates rheological modeling of EVP materials by using a physics-informed neural network (PINN) approach that embeds a modified Saramito model into the training framework. By directly fitting time-dependent stress data on typical EVP materials, this method circumvents the need for gradient estimation from noisy measurements and offers a data-efficient, differentiable, and generalizable alternative to conventional fitting methods. A shear-thinning formulation is introduced to reflect the decreasing viscosity at higher strain amplitudes. The proposed PINN approach is validated using both synthetic and experimental data, demonstrating stable recovery of physical parameters, enhanced interpretability, and improved predictive capability across a range of strain amplitudes. This framework bridges microstructural deformation modes with rheological modeling, offering a powerful tool for understanding and predicting nonlinear viscoelastic behavior in soft matter.
- Research Article
- 10.1016/j.neunet.2026.109251
- Jun 9, 2026
- Neural networks : the official journal of the International Neural Network Society
- Chengmao Wu + 1 more
Robust jointly sparse semi-supervised fuzzy C-means clustering with asymmetric deviation constraints.
- Research Article
- 10.1038/s41598-026-55768-3
- Jun 8, 2026
- Scientific reports
- Faisal S Alsubaei + 2 more
The rapid expansion of Internet of Things (IoT) infrastructures has significantly increased the exposure of edge devices to malware and botnet attacks. Conventional intrusion detection systems are largely centralized and struggle to operate effectively in decentralized, heterogeneous, and privacy-sensitive IoT environments, thereby limiting scalability and robustness. To address these challenges, this study proposes the Federated ConvNeXt-Swin Temporal Fusion Network (F-CSTFNet), a federated deep learning framework designed for distributed IoT malware and botnet detection. The proposed architecture integrates ConvNeXt-based convolutional feature extraction with Swin Transformer temporal attention to capture both local traffic patterns and long-range behavioral dependencies within network flows. This hybrid convolution-attention design enables the detection of short-term anomalies as well as evolving attack dynamics directly from network telemetry. In addition, a channel-adaptive feature recalibration mechanism enhances robustness when learning from heterogeneous and noisy client data. The model is trained using a federated learning paradigm that enables multiple IoT clients to collaboratively learn a global model without sharing raw data, thereby preserving data privacy and locality. Extensive experiments conducted on the IoT-23 and N-BaIoT datasets demonstrate that F-CSTFNet outperforms several state-of-the-art centralized and federated baselines in terms of detection accuracy, convergence stability, and client-level fairness. The framework also achieves low performance variance across clients, a high Jain's Fairness Index (JFI), and reduced inequality during distributed training. These results demonstrate the effectiveness of the proposed architecture as a scalable, privacy-preserving, and resilient intrusion detection framework for next-generation IoT security systems.
- Research Article
- 10.1038/s41598-026-55727-y
- Jun 7, 2026
- Scientific reports
- Fakir Mashuque Alamgir + 7 more
Real-time cyber-attack intrusion detection faces serious challenges in the smart grid communications infrastructure, as intrusion tactics become more advanced. Traditional rule-driven detection methods are unable to adapt to diverse attack patterns in modern power networks. In this work, a supervised deep learning framework is developed, using a CNN for spatial feature extraction, a BiLSTM for temporal dependency modeling, and an Extra Trees ensemble classifier to produce robust decisions, to achieve real-time intrusion detection on high-frequency smart meter data. The CNN layer extracts hierarchical spatial features from 128-dimensional multi-modal meter measurements (e.g., voltage, current, frequency harmonics), and the BiLSTM component captures temporal dynamics by processing whole sequences of meter data in both forward and backward directions to capture attack-evolution patterns that unidirectional models miss. Attention mechanisms dynamically weight the relevance of temporal features and enhance both prediction accuracy and interpretability. The Extra Trees ensemble provides a robust, low-variance decision output as an alternative to a standard Softmax layer. The architecture addresses several challenges: class imbalance (2.49:1 ratio), high dimensionality and noisy sensor data from heterogeneous sources, the requirement for real-time (millisecond-level) inference, and the need for explainability. The model was evaluated on 72,073 labeled power-grid logs from the Mississippi State University Power Grid Testbed, including False Data Injection Attacks, denial-of-service, replay, and man-in-the-middle attacks versus normal operation. With stratified five-fold cross-validation, the model achieves accuracy of 92.17% (± 0.27%), precision of 90.58% (± 0.49%), recall of 81.24% (± 0.86%), F1-score of 85.66% (± 0.20%), and ROC-AUC of 95.60% (± 0.14%), with an inference latency of only 12ms, which is suitable for utility-scale deployment. A comprehensive comparative study against nine imbalance-handling strategies (no handling, class weighting, SMOTE, ADASYN, Borderline-SMOTE, SMOTETomek, SMOTEENN, random over/under-sampling) confirms the chosen weighted-learning strategy as a Pareto-optimal choice for this dataset. Paired McNemar's tests with Holm correction demonstrate that the proposed model's improvements over the tested baselines are statistically significant ([Formula: see text]). Ablation studies validate that bidirectional processing increases accuracy by 20.12pp over a unidirectional CNN-LSTM, and that the Extra Trees head boosts precision by 10.09pp over the standalone Extra Trees baseline. This work contributes to hybrid deep learning for cyber-physical systems and provides deployment guidance in terms of computational cost and a human-in-the-loop framework. In future work, we will investigate cross-dataset validation, multi-simultaneous attack detection, graph neural networks for fault location, and federated learning toward privacy-preserving collaborative training.
- Research Article
- 10.3390/pr14121846
- Jun 7, 2026
- Processes
- Changyong Li + 7 more
Long horizontal wells in high-permeability fault-block reservoirs may intersect multiple faults, leading to complex pressure-transient responses, strong parameter coupling in conventional well-test interpretation, inefficient manual history matching, and pronounced non-uniqueness in fault-property identification. To address these challenges, this study proposes a physics-regularized neural inversion framework based on a PINN parameterization and low-weight physics regularization for well-test parameter inversion in long horizontal wells intersecting multiple faults. The proposed method takes the multiple-fault pressure response of a long horizontal well as the target problem. Both the pressure–drawdown curve and the pressure–drawdown derivative curve are used as data constraints. At the same time, parameter scaling and stage-wise training are introduced to jointly invert the reservoir permeability, fault transmissibility coefficient, skin factor, and effective producing length of the horizontal well. Considering that the simplified line-source forward model is not fully consistent with the two-dimensional pressure-diffusion equation and the fault-interface residuals, a physics-loss consistency test is performed to determine safe weighting ranges for the PDE residual and the fault-interface residual. These residuals are then incorporated into the training process as low-weight physics regularization terms to improve the physical plausibility of the inversion results. Results from the base case, different fault types, multiple-fault combinations, noise-robustness tests, ablation experiments, and method comparisons show that the proposed method can stably fit pressure–drawdown and pressure–drawdown derivative curves and effectively identify key well-test parameters in single-fault cases and some multiple-fault cases. In single-fault cases, the order of magnitude of the fault transmissibility coefficient can be identified stably. Reliable inversion performance is obtained for medium- to high-transmissibility faults and some multiple-fault combinations. In contrast, ambiguity remains between sealing faults and strong-baffle faults in multiple low-transmissibility fault combinations. The results further indicate that, under multiple random initializations, the physics-regularized neural inversion framework provides improved inversion stability in the tested synthetic low-transmissibility multiple-fault cases compared with the traditional least-squares method. Therefore, the proposed framework can serve as an intelligent auxiliary tool for well-test parameter inversion and fault-connectivity evaluation in complex fault-block reservoirs. Nevertheless, fine discrimination of low-transmissibility faults and interpretation of highly noisy field data still require joint constraints from geological, seismic, and production-dynamic information. A preliminary reduced field PINN fitting test using the well X falloff event further provides an engineering-scale applicability check for real pressure-transient data, with a pressure NRMSE of 2.457% for the extracted shut-in response.
- Research Article
- 10.1038/s41598-026-56128-x
- Jun 5, 2026
- Scientific reports
- Wang Junce
Data-driven approaches of instructional design are more necessary than ever for higher education institutions in the era of digital transformation, especially when faced with noisy and heterogenous educational data. In the digital transformation age, the requirement for data-driven approaches to instructional design that are stable for educational data is growing. Past studies of smart education and problem-based learning (PBL) have been mostly descriptive models or single-dataset studies, with few studies providing evidence for robustness-aware instructional design under different data characteristics. To fill this gap, this study presents a robustness-aware multi-objective optimization method called Cross-Dataset Robust Multi-Objective Optimization (CD-RMO) for exploring PBL design configurations in intelligent learning environments. CD-RMO presents learning gain, development of digital literacy, and implementation cost/risk as three objectives to be optimized simultaneously, and adds robustness to data variability to the objective measurement. The framework is tested on three different heterogeneous datasets related to education: OULAD, EdNet, and KDD Cup 2010, by employing a common feature space and a common PBL design variable. Convergence, diversity, robustness and stability are used as indicators for performance assessment. Under the proposed modeling framework, the results show that CD-RMO yields higher quality Pareto-front, convergence accuracy and robustness consistency compared to the baselines based on experts, classical evolutionary algorithms and a state-of-the-art multi-objective optimizer. In addition, sensitivity and interaction analyses indicate that Scaffolding Strength, PBL Intensity, and Feedback Frequency are significant on robustness-aware optimization performance and other design variables have their effect mostly through interaction effects. Overall, the proposed framework shows that robustness-aware multi-objective optimization is a principled computational methodology to study various trade-offs in ID for intelligent higher education, offering a much more stable and interpretable framework for making decisions in the setting of realistic data uncertainty.