Articles published on Synthetic data
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
35891 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.neuroimage.2026.121985
- Jul 15, 2026
- NeuroImage
- Jingying Yang + 12 more
Rapid multi-parametric quantitative MRI via deep learning-based synthetic-to-real reconstruction and 3D SSFP-MOLED imaging.
- New
- Research Article
- 10.1109/tbme.2025.3635264
- Jul 1, 2026
- IEEE transactions on bio-medical engineering
- Francesco Prendin + 2 more
Accurately predicting glucose levels is essential for effectively managing type 1 diabetes (T1D), a chronic condition in which the body cannot produce insulin. Although deep learning approaches have shown promise, their training requires extensive datasets that capture a wide range of physiological and behavioral variations. However, obtaining such datasets can be challenging and impractical, especially when their collection demands significant patient effort. To overcome this limitation, we propose a data augmentation strategy that leverages digital twins of individuals with T1D (DT-T1D) to generate personalized synthetic data mirroring real-world glucose-insulin dynamics. ReplayBG, an open-source tool for creating DT-T1D, was adapted to develop a two-steps strategy: first, generating DT-T1D from retrospective patient data; then, using DT-T1D with new inputs, to simulate synthetic, patient-specific data. The practical impact of this approach is demonstrated in a case study where personalized deep networks were developed to predict glucose levels. Models were trained on an open-source dataset from 12 patients, using either the original data or a combination of the original and synthetic data. Integrating synthetic data into the training process consistently enhances model performance. Moreover, models trained on synthetic data combined with only a small fraction of the original dataset achieve results comparable to those obtained from the full, unaugmented dataset. Leveraging DT-T1D to generate personalized synthetic data mitigates data scarcity and enhances deep learning model performance for accurate glucose prediction. This work highlights the potential of digital twin-driven data augmentation to tackle data scarcity and develop robust, personalized predictive models for T1D management.
- New
- Research Article
- 10.1109/tpami.2026.3672655
- Jul 1, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Xun Yang + 5 more
One-shot Federated Learning (OFL) has emerged as a promising paradigm, enabling global model training with minimal communication overhead. In OFL, the server model is usually distilled from an ensemble of pre-trained client models, while the ensemble also facilitates synthetic data generation for the knowledge distillation process. Prior works show that the performance of the final model is fundamentally tied to both the quality of the synthetic data and the ensemble. However, existing methods often optimize these two components separately, overlooking their interaction. To address this coupled optimization problem and provide a unified solution to the dual challenges of data and model heterogeneity inherent in OFL, we introduce Co-Boosting++, a novel OFL framework where synthetic data generation and ensemble construction mutually enhance each other in an iterative fashion. First, we fix the ensemble and generate hard samples in an adversarial manner. These samples are crucial for enhancing the robustness of knowledge transfer, as they challenge the model to generalize better, thereby improving quality of the synthetic data and subsequent distillation process. Second, leveraging these hard samples, we enhance the ensemble via a Mixture of Experts (MoE) mechanism. MoE allows dynamic adjustment of ensemble weights based on the generated hard samples, which enables the ensemble to better capture diverse and heterogeneous knowledge from client models. Furthermore, we extend Co-Boosting++ to support the simultaneous generation of multiple heterogeneous target models, enabling efficient adaptation to diverse device constraints. Extensive experiments on benchmark datasets demonstrate that Co-Boosting++ consistently outperforms state-of-the-art methods due to its coupled optimization of data and ensemble quality. Additionally, Co-Boosting++ is highly practical in real-world model market scenarios, requiring no local training modifications, additional transmissions, or restrictions on client model architectures.
- New
- Research Article
- 10.57264/cer-2026-0074
- Jul 1, 2026
- Journal of comparative effectiveness research
- Paul Arora + 1 more
In this update, we review a framework for identifying and mitigating information bias in electronic health records and administrative claims data, highlighting practical recommendations for study design, variable definition, and statistical analysis. We also discuss a perspective on emerging privacy-preserving technologies - synthetic data and federated networks - that enable secure cross-border data access while maintaining patient privacy.
- New
- Research Article
- 10.1016/j.canlet.2026.218493
- Jul 1, 2026
- Cancer letters
- Ruichong Lin + 6 more
Artificial intelligence in clinical oncology: Multimodal integration and translational development.
- New
- Research Article
- 10.1109/tpami.2026.3669002
- Jul 1, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Lucas Nunes + 3 more
Semantic scene understanding is crucial for robotics and computer vision applications. In autonomous driving, 3D semantic segmentation plays an important role for enabling safe navigation. Despite significant advances in the field, the complexity of collecting and annotating 3D data is a bottleneck in this developments. To overcome that data annotation limitation, synthetic simulated data has been used to generate annotated data on demand. There is still, however, a domain gap between real and simulated data. More recently, diffusion models have been in the spotlight, enabling close-to-real data synthesis. Those generative models have been recently applied to the 3D data domain for generating scene-scale data with semantic annotations. Still, those methods either rely on image projection or decoupled models trained with different resolutions in a coarse-to-fine manner. Such intermediary representations impact the generated data quality due to errors added in those transformations. In this work, we propose a novel approach able to generate 3D semantic scene-scale data without relying on any projection or decoupled trained multi-resolution models, achieving more realistic semantic scene data generation compared to previous state-of-the-art methods. Besides improving 3D semantic scene-scale data synthesis, we thoroughly evaluate the use of the synthetic scene samples as labeled data to train a semantic segmentation network. In our experiments, we show that using the synthetic annotated data generated by our method as training data together with the real semantic segmentation labels, leads to an improvement in the semantic segmentation model performance. Our results show the potential of generated scene-scale point clouds to generate more training data to extend existing datasets, reducing the data annotation effort.
- New
- Research Article
- 10.1107/s2053273326004006
- Jul 1, 2026
- Acta crystallographica. Section A, Foundations and advances
- Hemant Sharma + 2 more
The inversion of large-scale diffraction datasets from modern synchrotron sources presents a fundamental challenge in computational crystallography. This paper presents a unified algorithmic framework for the analysis of both near-field (morphological) and far-field (orientational and strain) high-energy diffraction microscopy (HEDM) data. We detail the mathematical formalisms and physical models that form the foundation of this methodology. Key aspects include a generalized model for detector distortion correction, robust algorithms for peak identification in noisy and overlapping patterns, a computationally efficient indexing formalism based on Friedel pair symmetry, and a decoupled iterative refinement scheme that exploits the differing sensitivities of position, orientation and lattice parameters to diffraction observables. We also describe the synergistic integration of near-field and far-field data streams, a critical feature of a truly comprehensive approach. The framework is validated in Part II of this series [Sharma et al. (2026). Acta Cryst. A82, https://doi.org/10.1107/S2053273326004018] using both experimental Ti-7 Al datasets and synthetic reconstructions with known ground truth, achieving orientation accuracy of ∼0.05° and position accuracy of ∼10 µm on experimental data, and a 190× improvement in lattice parameter precision over conventional simultaneous parameter refinement on synthetic data. This integrated framework provides a powerful and extensible solution for turning raw diffraction images into actionable microstructural and micromechanical information.
- New
- Research Article
- 10.1016/j.neunet.2026.108733
- Jul 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Zheshun Wu + 5 more
Enhancing progressive ensemble learning via normalized extra-Gradient initialization.
- New
- Research Article
- 10.1016/j.autcon.2026.106924
- Jul 1, 2026
- Automation in Construction
- Wanru Yang + 5 more
Structural health monitoring of underground tunnels increasingly uses advanced sensing and data-driven methods. Laser-scanned 3D point clouds capture spatially rich measurements of segmental tunnel linings and require segmentation as a prerequisite for downstream analysis. Deep learning (DL) is effective for point-cloud segmentation, but scarce datasets and costly annotation limit practical use. This paper presents Tunnel Scanner , a high-fidelity simulator that synthesises realistic tunnel point clouds with automatic annotation. The plug-and-play module Hybrid Position–Normal Local Spatial Encoding embeds geometric priors into DL backbones and combines with transfer learning (TL) to exploit synthetic data for domain adaptation. Models trained only on synthetic data achieved 71.9% mean Intersection-over-Union (mIoU) and 86.6% Overall Accuracy (OA), and TL increased performance to at least 78.8% mIoU and 90.9% OA with limited real data. This paper highlights geometry-informed data synthesis as a viable augmentation approach for digital inspection and asset management of large-scale tunnels. • Develop a high-fidelity simulator to address the scarcity of tunnel point clouds. • Propose a plug-and-play module encoding geometric features into DL backbones. • Investigate transfer learning to enhance 3D Sim-to-Real domain adaptation. • The end-to-end framework yields +24% segmentation accuracy with limited real data.
- New
- Research Article
- 10.1016/j.eswa.2026.132154
- Jul 1, 2026
- Expert Systems with Applications
- David De-Fitero-Dominguez + 2 more
Self-bootstrapping automated program repair: using LLMs to generate and evaluate synthetic training data for bug repair
- New
- Research Article
- 10.1109/tvcg.2026.3677109
- Jul 1, 2026
- IEEE transactions on visualization and computer graphics
- Chenhao Guo + 4 more
Portrait relighting shows great potential in photography, film, and AR by simulating diverse lighting effects. Existing state-of-the-art methods often rely on expensive paired OLAT or synthetic data, which limits scalability. Moreover, accurately modeling the interaction between physics-guided rendering, neural rendering, and real-world remains challenging. To address these issues, we propose a novel multi-stage self-supervised relighting framework. It progressively refines intrinsic scene properties via a simple-to-complex training strategy, removing the need for expensive paired data while adapting to various lighting conditions. One core design introduces a novel pre-training method approach using diverse shading-based masking for self-reconstruction, which improves the model's perception of complex lighting variations. Furthermore, we introduce two perceptual modules that leverage the linear superposition of light to narrow the gap between physics-guided and neural rendering, and better align relit results with real-world observations. Extensive experiments demonstrate that our unified framework achieves new state-of-the-art performance in portrait relighting, surpassing recent methods in photorealism, synthesis quality, and identity preservation. It provides a practical paradigm for high-fidelity relighting under diverse lighting.
- New
- Research Article
- 10.1016/j.jbi.2026.105043
- Jul 1, 2026
- Journal of biomedical informatics
- Muhammad Aslanimoghanloo + 2 more
Generative modeling of clinical time series via latent stochastic differential equations.
- New
- Research Article
- 10.1016/j.ijthermalsci.2026.110805
- Jul 1, 2026
- International Journal of Thermal Sciences
- Isabela Florindo Pinheiro + 4 more
This study combines experimental and theoretical approaches to investigate steady-state, multi-dimensional heat conduction in polymeric fins. Surface temperature fields are measured using infrared thermography, while a normalized two-dimensional heat conduction model is developed and solved via integral transform techniques. Two base boundary conditions — prescribed temperature (Dirichlet) and prescribed heat flux (Neumann) — are analyzed to evaluate their impact on thermal behavior and parameter estimation. The Biot number is estimated using two approaches: (1) an inverse problem solved with the Levenberg–Marquardt (LM) algorithm and (2) a machine learning model based on Gradient Boosted Trees (GBT). Synthetic data generated from the forward model serve as the training set for the GBT approach, aligning with Problem-Informed Machine Learning (PIML) methodologies. Both approaches are then employed to estimate the convective heat transfer coefficient, while the material’s thermal conductivity is experimentally measured using a Heat Flow Meter (FOX 50). Results indicate that the LM method provides interpretability and strong performance when sensitivity is adequate and regularization is applied, while the GBT demonstrates greater robustness in nonlinear regimes and with ample training data. • Biot number estimation in polymeric fins via inverse problem and machine learning. • Synthetic temperature data from the 2D heat model trained ML for fast Bi prediction. • GBT achieved the best overall performance; Levenberg–Marquardt remained stable. • Infrared imaging experiments confirmed model validity and boundary assumptions. • Both methods achieved < 10 % average Biot error, validating the proposed approach.
- New
- Research Article
- 10.1016/j.neunet.2026.108727
- Jul 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Yufan Liao + 3 more
Decorr: Environment partitioning for invariant learning and OOD generalization.
- New
- Research Article
- 10.1016/j.socnet.2026.02.001
- Jul 1, 2026
- Social Networks
- Martina Boschi + 2 more
Technological advances enable the collection of many types of complex, dynamic, relational hyperevents. Despite the complexity of the data, most current Relational Hyper Event Models (RHEMs) use simple linear effects to describe the event rates. We extend the RHEM in order to allow non-linear effects that vary over time by using tensor product smooths. We validate our method on both synthetic and empirical data, examining evolving patterns and the impact of scientific collaboration between multiple players. Our approach provides new insights into hyperevent dynamics, uncovering potential non-monotonic patterns that linear models cannot capture. • Interpretation of linear RHEMs can dramatically fail in presence of non-monotonic effects. • No prior assumptions of linear time-homogeneous effects of drivers on event rate. • Inference in relational hyper event models framed as logistic additive regression. • Joint time-varying non-linear effects can be modeled as tensor product smooths. • Author self-citation, a scientific innovation driver, shows a non-monotonic effect.
- New
- Research Article
- 10.1109/tpami.2026.3673525
- Jul 1, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Banglei Guan + 2 more
In recent years, affine correspondences (ACs) have emerged as widely adopted alternative to point correspondences (PCs) in geometric problems in computer vision. An AC is composed of a PC across two different views plus an affine transformation between the small patches around this PC. Prior studies have shown that a single affine correspondence (AC) generally yields three independent constraints for estimating relative pose. This work addresses relative pose estimation in multi-perspective camera systems, a relevant problem given their prevalence in modern technologies such as autonomous vehicles and augmented reality. More specifically, we introduce the first comprehensive suite of minimal solvers for 6DoF relative pose estimation across multiple cameras using only two ACs, which is notably valuable for robust model fitting scenarios. We analyze all possible configurations of two ACs in two views, and present minimal solvers covering all identified minimal cases. We make use of the hidden variable technique to eliminate the translation parameters, and represent rotation using either Cayley parameters or quaternions. We furthermore introduce novel constraints on the generalized relative pose problem that are beneficial in deriving more compact solvers with fewer solutions. Comprehensive experiments on synthetic and real-world data show that the proposed affine correspondence-based solvers are highly effective and computationally efficient.
- New
- Research Article
- 10.1016/j.mimet.2026.107529
- Jul 1, 2026
- Journal of microbiological methods
- Abbas Karimi-Fard + 1 more
AC-WGAN-GP for transcriptomic data augmentation: Enhancing stress classification in Synechocystis sp. PCC 6803 under data scarcity.
- New
- Research Article
- 10.1016/j.bspc.2026.109934
- Jul 1, 2026
- Biomedical Signal Processing and Control
- S Janifer Jabin Jui + 5 more
Stress is a widespread concern that impacts human health with its silent progression, causing significant public health burdens and economic loss globally. Non-invasive wearable technology empowered by physiological signal monitoring can enable early warning systems for stress, alleviating some of the burdens, allowing on-time interventions, and thus significantly improving quality of life. This study used the heart rate and respiratory rate data from 34 participants. It evaluated the performance of the hybrid deep learning CNN-Transformer model and benchmarked it against deep learning convolutional neural networks (CNNs) and Transformer models, extreme gradient boosting (XGBoost) and random forest (RF) machine learning models, comprising a total of five AI models. To mitigate data imbalance and observe the efficacy of deep learning data augmentation techniques in physiological signals for stress monitoring, two generative adversarial network (GAN) models: conditional tabular GAN (CTGAN), copula GAN (CopGAN) and variational autoencoder (VAE) based model tabular VAE synthesiser (TVAES) had been employed. The modelling performance significantly improved when applying CTGAN and CopGAN, demonstrating the usefulness of synthetic data. The CNN-Transformer achieved an average accuracy of 77%, a precision of 87% and an AUC of 83%. The study applied leave-one- subject- out (LOSO CV) to prove the CNN-Transformer hybrid’s robustness for generalizability to perform well for unseen subjects. The study integrated explainable AI models, Shapley values (SHAP), and local interpretable model-agnostic explanations (LIME), as well as Monte Carlo Dropout for uncertainty quantification to bring confidence, trust and transparency to AI systems, taking a step closer to real-world deployment. Similar studies can also help in the detection of other disorders, such as anxiety and depression. • Comprehensive multi-domain feature extraction from heart rate and respiratory rate using wearable non-invasive sensor technology. • Comparative evaluation of three deep learning data augmentation techniques (CTGAN, TVAES and CopulaGAN) to mitigate data imbalance. • Comparative performance analysis of the hybrid DL model, CNN-Transformer model against benchmarking DL and ML models. • Implementing uncertainty quantification using Monte Carlo Dropout for confident and reliable model prediction assessment. • Application of xAI models LIME and SHAP for interpretability and feature importance.
- New
- Research Article
- 10.1016/j.optcom.2026.132993
- Jul 1, 2026
- Optics Communications
- Jianan Fu + 5 more
Scattering imaging beyond the optical memory effect using deep learning model trained on physically inspired synthetic data
- New
- Research Article
- 10.1016/j.cmpb.2026.109352
- Jul 1, 2026
- Computer methods and programs in biomedicine
- Jakub Gumulski + 3 more
Semi-automatic generation of selected cerebral vessels for the objective evaluation of vessel segmentation and their geometric parameters in computed tomography angiography images.