Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

An open benchmark dataset of synthetic seismic data and real swell noise for evaluating deep learning denoising models

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

An open benchmark dataset of synthetic seismic data and real swell noise for evaluating deep learning denoising models

Similar Papers
  • Research Article
  • Cite Count Icon 7
  • 10.1190/tle41060392.1
Synthetic seismic data generation for automated AI-based procedures with an example application to high-resolution interpretation
  • Jun 1, 2022
  • The Leading Edge
  • Fernando Vizeu + 5 more

This paper discusses the generation of synthetic 3D seismic data for training neural networks to solve a variety of seismic processing, interpretation, and inversion tasks. Using synthetic data is a way to address the shortage of seismic data, which are required for solving problems with machine learning techniques. Synthetic data are built via a simulation process that is based on a mathematical representation of the physics of the problem. In other words, using synthetic data is an indirect way to teach neural networks about the physics of the problem. An important incentive for using synthetic data to solve problems with artificial intelligence methods is that with real seismic data the ground truth is always unknown. When generating synthetic seismic data, we first build the model and then calculate the data, so the answer (model) is always known and always exact. We describe a methodology for generating on-the-fly simulated postmigration (1D modeling) synthetic data in 3D, which are high resolution and look similar to real data. A wide range of models is covered by generating an unlimited number of data examples. The synthetic data are built from impedance models that are constructed through geostatistical simulation of real well logs. With geostatistical simulation, we can describe various geologic variance models in 3D and obtain realistic images. To cover a broad range of scenarios, we need to generalize the seismic data story by randomly perturbing many parameters including structures, conformity styles, dip-strike directions, variograms, measured input logs, frequencies, phase spectra, etc.

  • Research Article
  • Cite Count Icon 10
  • 10.1190/int-2021-0193.1
Synthetic seismic data for training deep learning networks
  • May 17, 2022
  • Interpretation
  • Tom P Merrifield + 7 more

Deep learning is increasingly being used as a component of geoscience workflows for processing and interpreting seismic data. Training a supervised deep learning network is a data-hungry task: a significant number of data examples are needed and they must include labels. The data examples and their labels must have consistent patterns for the deep learning network to learn. Too few examples and/or poor-quality labels can lead to poor deep learning training results. One method to provide large quantities of training examples with high-quality labels is to create synthetic data. We discuss our techniques and experiences with our ongoing use of synthetic seismic data. We share our techniques as an open-source project concurrent with this paper at https://github.com/tpmerrifield/synthoseis. We hope that the geoscience community will share our enthusiasm for developing deep learning geoscience tools and for including synthetic seismic data in supervised deep learning training. We invite contributions from the geoscience community using the open-source model to collectively reduce the realism gap between synthetic data and field seismic data.

  • Research Article
  • Cite Count Icon 20
  • 10.1109/tgrs.2022.3213337
Seismic Inversion Based on 2D-CNNs and Domain Adaption
  • Jan 1, 2022
  • IEEE Transactions on Geoscience and Remote Sensing
  • Qi Wang + 3 more

Deep learning has been applied to tackle the seismic inversion problem, bringing more efficiency and accuracy. However, bad spatial continuity and poor generalizability limit the practical application. To solve these problems, we propose a 2D end-to-end seismic inversion method based on domain adaption. Firstly, the proposed 2D network learns the inversion mapping of seismic data under the constraint of domain adaption layer, which can reduce the difference between the features of real seismic data and synthetic seismic data, improving the generalization ability on real seismic data. Then, the trained model is finetuned with well logging data. In the first process, the spatial continuity of the inversion result is guaranteed by the 2D training scheme. Meanwhile, due to the constraint of the domain adaption layer, our model not only performs well on the synthetic data but also has good generalization ability on the real seismic data. And we carefully discuss the mechanism of domain adaption layer. In the second process, finetuning introduces well logging information, which can further improve the ability to invert details. Moreover, in order to improve the inversion accuracy on real seismic data, we develop a new training data generation method that can generate the synthetic samples close to the real samples, and a 2.5D training strategy is adopted to improve the continuity of the 3D data. The experiments on both synthetic and real seismic data show that our method performs better than both the recursive inversion method and the 1D closed-loop CNN methods.

  • Research Article
  • Cite Count Icon 1
  • 10.1016/j.aiig.2025.100109
Robust low frequency seismic bandwidth extension with a U-net and synthetic training data
  • Jun 1, 2025
  • Artificial Intelligence in Geosciences
  • P Zwartjes + 1 more

Robust low frequency seismic bandwidth extension with a U-net and synthetic training data

  • Conference Article
  • Cite Count Icon 3
  • 10.20396/revpibic2720192342
Enriching synthetic data with real noise using Neural Style Transfer
  • Nov 30, 2019
  • Resumos do...
  • Naomi Takemoto + 5 more

Deep Learning experiments require large amounts of labeled data, but few annotated seismic datasets are available and annotation is a time-consuming, expensive activity. Synthetic modeled datasets may be a viable alternative. However, they lack the variability and intricacies of a real data signal. Moreover, methods that add colored noises are not enough to represent such variability. Thus, our goal is to produce a noise type that is characteristic of real data to create a viable synthetic dataset to train Deep Learning models. In this context, we apply the Neural StyleTransfer technique, which combines the structural content of an image with the textural style of another, to produce a synthetic data with noise characteristics extracted from a real dataset. The results show that the stylized synthetic seismic data preserves the modeled content while incorporating characteristics of some real data chosen as style, creating synthetic data with a more realistic noise profile.

  • Research Article
  • Cite Count Icon 9
  • 10.1190/geo2021-0568.1
Consistency and prior falsification of training data in seismic deep learning: Application to offshore deltaic reservoir characterization
  • Apr 11, 2022
  • GEOPHYSICS
  • Anshuman Pradhan + 1 more

Deep learning (DL) applications of seismic reservoir characterization often require the generation of synthetic data to augment available sparse labeled data. An approach for generating synthetic training data consists of specifying probability distributions modeling prior geologic uncertainty on reservoir properties and forward modeling the seismic data. A prior falsification approach is critical to establish the consistency of the synthetic training data distribution with real seismic data. With the help of a real case study of facies classification with convolutional neural networks (CNNs) from an offshore deltaic reservoir, we have highlighted several practical nuances associated with training DL models on synthetic seismic data. We highlight the issue of overfitting of CNNs to the synthetic training data distribution and propose regularization strategies to address it. We demonstrate the efficacy of our proposed strategies by training the CNN on synthetic data and making robust predictions with real 3D partial stack seismic data.

  • Research Article
  • Cite Count Icon 63
  • 10.1016/j.aiig.2022.09.002
MLReal: Bridging the gap between training on synthetic data and real data applications in machine learning
  • Nov 7, 2022
  • Artificial Intelligence in Geosciences
  • Tariq Alkhalifah + 2 more

MLReal: Bridging the gap between training on synthetic data and real data applications in machine learning

  • Research Article
  • Cite Count Icon 39
  • 10.1111/1365-2478.13097
Learning from unlabelled real seismic data: Fault detection based on transfer learning
  • Jun 6, 2021
  • Geophysical Prospecting
  • Ruoshui Zhou + 3 more

ABSTRACTSignificant advances have been made towards fault detection using deep learning. However, the fault labelling of seismic data requires great human effort. The resulting small sample problem makes traditional deep learning methods difficult to achieve desired results. Existing research proposes to train a deep learning model with labelled synthetic seismic data to get good fault detection results. However, due to the complexity of the actual geological situation, there are inevitable differences between synthetic seismic data and real seismic data in many aspects such as seismic signal frequency, frequency of fault distribution and degree of noise disturbance, which lead to the fact that the deep learning model trained by synthetic seismic data is difficult to get good fault detection result in field data applications. We propose to use transfer learning to reduce the impact of data differences to solve this problem: part of the deep transfer learning model is used to learn fault‐related features. And the other part of the deep transfer learning model is used to mine common features between the real seismic data and the synthetic seismic data, which makes the deep transfer learning model more suitable for real seismic data. Compared with the latest research progress, our method can greatly improve the effect of fault detection without real data label, which can significantly save the cost of manual label processing.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.2516/ogst/2011168
Monitoring of SAGD Process: Seismic Interpretation of Ray+Born Synthetic 4D Data
  • Mar 1, 2012
  • Oil & Gas Science and Technology – Revue d’IFP Energies nouvelles
  • C Joseph + 6 more

The objective of this study is to evaluate which production information can be deduced from a 4D seismic survey during the Steam-Assisted Gravity Drainage (SAGD) recovery process. Superimposed on reservoir heterogeneities of geological origin, many factors interact during thermal production of heavy oil and bitumen reservoirs, which complicate the interpretation of 4D seismic data: changes in oil viscosity, in fluid saturations, in pore pressure and so on. This study is based on the real Hangingstone field case of the McMurray formation in the Athabasca region (Canada). In previous works, an initial static model (geology, petroacoustic and geomechanical) has been constructed and a thermal production of heavy oil with two coupled fluid-flow and geomechanical models has been simulated. Seismic parameters (density, compression velocity and shear velocity) of the saturated rocks have then been computed from mechanical and reservoir parameters at several stages of the production. A repeated acquisition survey is modelled at different stages of SAGD production. This is performed using a 3D seismic modelling approach. To focus on the reflections generated within the reservoir zone, a target-oriented modelling is chosen. It is based on the ray+Born approach which permits to compute the P-wave elastic response by correctly handling the seismic amplitudes as a function of source-receiver offset. Real incoherent noise is added to the zero-phase synthetic data to produce a more realistic result. The noise-free and the noisy synthetic data are processed to get stacked and time migrated images. A simple processing workflow leads to image the steam chamber development, in particular its V-shape in radial section, and to observe time-lapse in the reservoir zone. An interpretation work is then carried out. Some seismic attributes like RMS values of amplitude changes between stages, energy, time differences of reservoir bottom between stages, etc. are computed from the synthetic (noise-free and noisy) seismic data. Some of these attributes prove to be robust to the noise and to show some production effect. Possible trends between these attributes and the modelled reservoir/geomechanical properties (lithofacies, pressure, temperature, steam saturation, etc.) are also evaluated. Finally, geobodies are extracted from the seismic attributes.

  • Research Article
  • Cite Count Icon 11
  • 10.1109/lgrs.2022.3218911
Deep Learning Using Synthetic Seismic Data by Fourier Domain Adaptation in Seismic Structure Interpretation
  • Jan 1, 2022
  • IEEE Geoscience and Remote Sensing Letters
  • Dekuan Chang + 5 more

Deep learning is a data-driven technique that demands network models trained using big datasets. In seismic structure interpretation, it is very difficult, time-consuming, and relatively economic costs to prepare training datasets by directly annotating real seismic data. Seismic convolution method is an efficient way to synthesize seismic data, which can easily and quickly generate large amount of training datasets. However, there are large differences in feature space between synthetic seismic data and real seismic data. That results in the poor performance of network models trained using synthetic training datasets on real seismic data. In this paper, we propose to use Fourier domain adaptation (FDA) to achieve domain transfer. First, amplitude spectrum of synthetic seismic data is replaced with those of real seismic data to make feature space mapping. Then synthetic seismic data is used for transfer training of network models to improve its performance on real seismic data. The experimental results demonstrate that the FDA performs the domain transfer of synthetic seismic data to real seismic data, which improves the generalization of network models trained based on synthetic seismic data. Meanwhile, the FDA is a promising method for deep learning model transfer training in seismic structure interpretation.

  • Research Article
  • Cite Count Icon 44
  • 10.1016/j.ijggc.2018.06.011
Modeling of time-lapse seismic monitoring using CO2 leakage simulations for a model CO2 storage site with realistic geology: Application in assessment of early leak-detection capabilities
  • Jun 28, 2018
  • International Journal of Greenhouse Gas Control
  • Zan Wang + 3 more

Modeling of time-lapse seismic monitoring using CO2 leakage simulations for a model CO2 storage site with realistic geology: Application in assessment of early leak-detection capabilities

  • Research Article
  • Cite Count Icon 7
  • 10.1002/nsg.12273
Adding realistic noise models to synthetic ground‐penetrating radar data
  • Oct 3, 2023
  • Near Surface Geophysics
  • Sophie Marie Stephan + 2 more

ABSTRACTCost‐effective computing capabilities have paved the road for the use of numerical modelling to develop advanced methods and applications of ground‐penetrating radar (GPR). Realistic synthetic data and the corresponding modelling techniques, respectively, should consider all subsurface and above‐ground aspects that influence GPR wave propagation and the characteristics of recorded signals. Critical aspects that can be realized in modern GPR modelling tools include heterogeneous and frequency‐dependent material properties, complex structures and interface geometries as well as three‐dimensional antenna models, including the interaction between the antenna and the subsurface. However, realistic noise related to the electronic components of a GPR system or ambient electromagnetic noise is often not considered, or simplified by assuming a white Gaussian noise model which is added to the modelled data. We present an approach to include realistic noise scenarios as typically observed in GPR field data into the flow of modelling synthetic GPR data. In our approach, we extract the noise from recorded GPR traces and add it to the modelled GPR data via a convolution‐based process. We illustrate our methodology using a modelling exercise, where we contaminate a synthetic two‐dimensional GPR dataset with frequency‐dependent noise recorded in an urban environment. Comparing our noise‐contaminated synthetic data with field data recorded in a similar environment illustrates that our method allows the generation of synthetic GPR with realistic noise characteristics and further highlights the limitations of assuming pure white Gaussian noise models.

  • Research Article
  • Cite Count Icon 9
  • 10.1109/tgrs.2023.3262884
Deep Learning With Fault Prior for 3-D Seismic Data Super-Resolution
  • Jan 1, 2023
  • IEEE Transactions on Geoscience and Remote Sensing
  • Ruoshui Zhou + 5 more

Seismic data are always in low-resolution due to the limitations of seismic acquisition and processing technology, which bring challenges to subsequent seismic interpretation. Deep learning has been successfully applied to the seismic data super-resolution, all methods tend to directly learn the mapping relationship between low-resolution seismic images and super-resolution seismic images through some complex convolutional neural networks. But blindly increasing the depth of the network brings limited improvement to the super-resolution work. We propose a novel geophysical prior guided framework for seismic data super-resolution to solve this problem. Specifically, we select fault prior to guide the training of deep learning model: we use knowledge distillation technology to progressively propagate the fault prior from the teacher network (trained with the low-resolution synthetic seismic data/high-resolution fault prior and high-resolution synthetic seismic data pairs) to the student net-work (trained with the low-resolution synthetic seismic data and high-resolution synthetic seismic data pairs). To better propagate fault priors, we use feature space loss and soft ground truth loss in student network training. Finally, we use the trained student network to complete the super-resolution of synthetic validation seismic data and real seismic data. In addition, in our super-resolution framework, we directly process 3D seismic data instead of 2D seismic images, which further improves the effects of super-resolution and the subsequent seismic interpretation. Compared with the state-of-the-art seismic super-resolution method, the experimental results show that the super-resolution results of our method can depict the faults more clearly.

  • Conference Article
  • 10.1190/segj102011-001.44
Partial CRS stack seismic data characterization on AVO anomaly: case study on 2D synthetic seismic data
  • Jan 1, 2011
  • Proceedings of the 10th SEGJ International Symposium, Kyoto, Japan, 20-22 November 2011
  • Alpius Dwi Guntara + 2 more

Partial CRS stack is one of alternative seismic data processing methods, which does not depend on macro velocity model. One of the advantages of this method compared with conventional CMP stack is the summation process in each CMPs in order to obtain better signal quality. CRS travel time formula can be applied to obtain partial stacked CRS supergathers. The CRS gathers obtained by this method are regularized on offset domain and have better S/N ratio than conventional gathers. This method is very helpful to improve the quality of land seismic data subject to the random noise and lack of fold coverage. The purpose of this research is to examine the efficiency of the partial CRS stack methods to highlight AVO/A (Amplitude Versus Offset/Angle) anomalies by using synthetic seismic data. The tested AVO anomaly is class 3 where the negative amplitude increases with increasing distance/offset. Seismic modeling is performed on a simple anticline model composed of shale and sand layers. The synthetic seismic data are provided to the processing using partial CRS stack to test the offset aperture that is considered to influence the seismic amplitude. AVO analysis was conducted on the pre-stack time migrated (PSTM) synthetic data with and without partial CRS stack for three offset apertures. The seismic amplitude responses change with the apertures. Dipping reflectors also give different amplitude response from flat reflectors. It can be inferred from the analysis that although the class of AVO anomalies on the partial CRS stack gathers are almost same, there should be an optimum offset aperture to avoid ambiguous result of the AVO attributes.

  • Preprint Article
  • 10.5194/egusphere-egu25-288
Forward Stratigraphic Modelling to Generate Synthetic Seismic Training Dataset for Deep Learning: A Case Study to Predict Shelf-Edge Trajectories
  • Mar 18, 2025
  • Waleed Algharbi + 2 more

Seismic data forms the backbone of what we understand in the subsurface, and seismic data interpretation is still usually done by hand. Automatic seismic interpretation with deep learning is very promising, but there the problem is a lack of labelled training data. In this study, we use forward stratigraphic modelling and show how forward modelling can be advantageously used in deep learning.Specifically, we focus on shelf-edge trajectories as the geological representations of lateral and vertical shifts in sediments’ position through time. They provide continuous tracks of changes in relative sea-level as well as sediment stacking patterns and depositional geometries. Mapping these trajectories and measuring their changing angles help in quantifying the sequence stratigraphic analysis and predicting ancient depositional environments.Here, we evaluate the ability of deep learning models, trained on synthetic seismic data, to identify clinoforms and their rollover points for shelf-edge trajectories mapping. The synthetic training dataset generated using geological processed-based forward modelling represents different depositional slope scenarios. Controlling the different parameters that govern shelf-edges and shelf-edge trajectories (such as bathymetry, sediment supply, eustatic sea-level changes and subsidence) gave us a better chance to mimic realistic and diverse depositional setting, which helps in generalizing the deep learning model. In addition, the ground truth (labels) for the created synthetic seismic data is automatically generated by the forward model, without the need of manual labelling seismic data.Higher accuracy score on both validation and testing datasets demonstrates the power and effectiveness of using synthetic as training dataset. This study shows that synthetic data can play a major role in bridging the gap between traditional seismic interpretation and automating the process using machine learning. It also shows that forward modelling is a powerful technique to combine with data modelling, such as machine learning.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant