Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Small angle neutron scattering in McStas: Optimization for high-throughput virtual experiments

  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

In this work, we present the development of small-angle scattering components in McStas that describe the neutron interaction with 70 different form and structure factors. We describe the considerations taken into account for the generation of these components, such as the incorporation of polydispersity and orientational distribution effects in the Monte Carlo simulation. These models can be parallelized by means of multi-core simulations and graphical processing units. The acceleration schemes for the aforementioned models are benchmarked, and the resulting performance is presented. This allows the estimation of computation times in high-throughput virtual experiments. The presented work enables the generation of large datasets of virtual experiments that can be explored and used by machine learning algorithms.

Similar Papers
  • Research Article
  • Cite Count Icon 15
  • 10.1002/cpe.1771
Pricing barrier and American options under the SABR model on the graphics processing unit
  • Jun 7, 2011
  • Concurrency and Computation: Practice and Experience
  • Yu Tian + 3 more

SUMMARYIn this paper, we presented our study on using the graphics processing unit (GPU) to accelerate the computation in pricing financial options. We first introduced the GPU programming and the SABR stochastic volatility model. We then discussed pricing options with quasi Monte Carlo techniques under the SABR model. In particular, we focused on pricing barrier options by quasi Monte Carlo and conditional probability correction methods and on pricing American options by the least squares Monte Carlo method. We then presented our GPU‐based implementation for pricing barrier options and hybrid CPU–GPU implementation for pricing American options. In addition, we described techniques for efficient use of GPU memory. We provided details of implementing these GPU numerical schemes for pricing options and compared performances of the GPU programs with their CPU counterparts. We found that GPU‐based computing schemes can achieve 134 times speedup for pricing barrier options, while maintaining satisfactory pricing accuracy. For pricing American options, we also reported that when the least squares Monte Carlo method is used, special techniques can be devised to use less GPU memory, resulting in 22 times speedup, instead of the original 10 times speedup. Copyright © 2011 John Wiley & Sons, Ltd.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 24
  • 10.3390/electronics13061010
Plant Disease Identification Using Machine Learning Algorithms on Single-Board Computers in IoT Environments
  • Mar 7, 2024
  • Electronics
  • George Routis + 2 more

This paper investigates the usage of machine learning (ML) algorithms on agricultural images with the aim of extracting information regarding the health of plants. More specifically, a custom convolutional neural network is trained on Google Colab using photos of healthy and unhealthy plants. The trained models are evaluated using various single-board computers (SBCs) that demonstrate different essential characteristics. Raspberry Pi 3 and Raspberry Pi 4 are the current mainstream SBCs that use their Central Processing Units (CPUs) for processing and are used for many applications for executing ML algorithms based on popular related libraries such as TensorFlow. NVIDIA Graphic Processing Units (GPUs) have a different rationale and base the execution of ML algorithms on a GPU that uses a different architecture than a CPU. GPUs can also implement high parallelization on the Compute Unified Device Architecture (CUDA) cores. Another current approach involves using a Tensor Processing Unit (TPU) processing unit carried by the Google Coral Dev TPU Board, which is an Application-Specific Integrated Circuit (ASIC) specialized for accelerating ML algorithms such as Convolutional Neural Networks (CNNs) via the usage of TensorFlow Lite. This study experiments with all of the above-mentioned devices and executes custom CNN models with the aim of identifying plant diseases. In this respect, several evaluation metrics are used, including knowledge extraction time, CPU utilization, Random Access Memory (RAM) usage, swap memory, temperature, current milli Amperes (mA), voltage (Volts), and power consumption milli Watts (mW).

  • Research Article
  • 10.1118/1.4735083
SU-E-T-28: Accelerated Event-By-Event Microdosimetry Monte Carlo Simulations of Low Energy Electron and Proton on a CUDA-Enabled GPU
  • Jun 1, 2012
  • Medical Physics
  • G Kalantzis + 3 more

Purpose: For microdosimetric calculation event-by-event Monte Carlo (MC) simulation is considered the most accurate, but it is very time-consuming. In this work we present an event-by-event MC simulation of low energy electron and proton for accelerated microdosimetric MC simulations on a graphic processing unit (GPU). Methods: The MC simulation of particle was implemented in C and executed on a multi-core CPU, and a commercially available general purpose GPU using the compute unified device architecture (CUDA). Additionally, a hybrid implementation scheme was realized by employing OpenMP and CUDA so that both GPU and multi-core CPU were utilized simultaneously. The two implementation schemes have been tested and compared with the sequential single threaded MC simulation on the CPU. A calculation time comparison was established on the speed-up for a set of benchmarking cases of electron and proton. Results: Dosimetric results were obtained with both the parallel and serial MC codes. A GPU over CPU speed-up of 67.2 and 19.2 times was achieved for 300 eV and 2 keV electron tracks respectively. For proton tracks, the GPU-based code was approximately 5 times faster than the CPU-single-thread code. By incorporating a multi-core CPU and running the MC code simultaneously on the GPU and CPU, an increase of 2%-7% and 20% in the speedup was noticed for electrons and protons tracks respectively. A good dosimetric agreement between the CPU- and GPU-based MC methods was observed for both electrons and proton. Conclusions: A GPU-based MC method for microdosimetric calculations of low energy electron and proton has been presented. The results indicate the capability of our GPU-based implementation for accelerated MC simulations of both electron and proton without loss of accuracy. Lastly, the potential of a hybrid approach by utilizing simultaneously a GPU and a multi-core CPU for further acceleration of MC microdosimetric calculations has been demonstrated.

  • Research Article
  • Cite Count Icon 1
  • 10.1088/1361-6560/adfda7
Review of GPU-based Monte Carlo simulation platforms for transmission and emission tomography in medicine
  • Aug 29, 2025
  • Physics in Medicine & Biology
  • Yujie Chi + 3 more

Objectives. Monte Carlo (MC) simulation remains the gold standard for modeling complex physical interactions in transmission and emission tomography, with graphic processing unit (GPU) parallel computing offering unmatched computational performance and enabling practical, large-scale MC applications. In recent years, rapid advancements in both GPU technologies and tomography techniques have been observed. Harnessing emerging GPU capabilities to accelerate MC simulation and strengthen its role in supporting the rapid growth of medical tomography has become an important topic. To provide useful insights, we conducted a comprehensive review of state-of-the-art GPU-accelerated MC simulations in tomography, highlighting current achievements and underdeveloped areas.Approach. We reviewed key technical developments across major tomography modalities, including computed tomography (CT), cone-beam CT (CBCT), positron emission tomography (PET), single-photon emission CT, proton CT , emerging techniques, and hybrid modalities. We examined MC simulation methods and major CPU-based MC platforms that have historically supported medical imaging development, followed by a review of GPU acceleration strategies, hardware evolutions, and leading GPU-based MC simulation packages. Future development directions were also discussed.Main results. Significant advancements have been achieved in both tomography and MC simulation technologies over the past half-century. The introduction of GPUs has enabled speedups often exceeding 100-1000 times over CPU implementations, providing essential support to the development of new imaging systems. Emerging GPU features like ray-tracing cores, tensor cores, and GPU-execution-friendly transport methods offer further opportunities for performance enhancement.Significance. GPU-based MC simulation is expected to remain essential in advancing medical emission and transmission tomography. With the emergence of new concepts such as training machine learning with synthetic data, digital twins for healthcare, and virtual clinical trials, improving hardware portability and modularizing GPU-based MC codes to adapt to these evolving simulation needs represent important directions for future research. This review aims to provide useful insights for researchers, developers, and practitioners in the relevant fields.

  • Conference Article
  • Cite Count Icon 2
  • 10.1115/gt2021-59295
High-Performance Computing Probabilistic Fracture Mechanics Implementation for Gas Turbine Rotor Disks on Distributed Architectures Including Graphics Processing Units (GPUs)
  • Jun 7, 2021
  • Mrugesh Gajjar + 2 more

We present an efficient Monte Carlo based probabilistic fracture mechanics simulation implementation for heterogeneous high-performance (HPC) architectures including CPUs and GPUs. The specific application focuses on large heavy-duty gas turbine rotor components for the energy sector. A reliable probabilistic risk quantification requires the simulation of millions to billions of Monte Carlo (MC) samples. We apply a modified Runge-Kutta algorithm in order to solve numerically the fatigue crack growth for this large number of cracks for varying initial crack sizes, locations, material and service conditions. This compute intensive simulation has already been demonstrated to perform efficiently and scalable on parallel and distributed HPC architectures including hundreds of CPUs utilizing the Message Passing Interface (MPI) paradigm. In this work, we go a step further and include GPUs in our parallelization strategy. We develop a load distribution scheme to share one or more GPUs on compute nodes distributed over a network. We detail technical challenges and solution strategies in performing the simulations on GPUs efficiently. We show that the key computation of the modified Runge-Kutta integration step speeds up over two orders of magnitude on a typical GPU compared to a single threaded CPU. This is supported by our use of GPU textures for efficient interpolation of multi-dimensional tables utilized in the implementation. We demonstrate weak and strong scaling of our GPU implementation, i.e., that we can efficiently utilize a large number of GPUs/CPUs in order to solve for more MC samples, or reduce the computational turn-around time, respectively. On seven different GPUs spanning four generations, the presented probabilistic fracture mechanics simulation tool ProbFM achieves a speed-up ranging from 16.4x to 47.4x compared to single threaded CPU implementation.

  • Conference Article
  • Cite Count Icon 10
  • 10.1109/nssmic.2012.6551511
Hybrid GATE: A GPU/CPU implementation for imaging and therapy applications
  • Oct 1, 2012
  • Julien Bert + 13 more

International audience

  • Book Chapter
  • Cite Count Icon 6
  • 10.1007/978-981-287-134-3_15
Soft Computing Methods for Big Data Problems
  • Sep 12, 2014
  • Shafaatunnur Hasan + 2 more

Generally, big data computing deals with massive and high-dimensional data such as DNA microarray data, financial data, medical imagery, satellite imagery, and hyperspectral imagery. Therefore, big data computing needs advanced technologies or methods to solve the issues of computational time to extract valuable information without information loss. In this context, generally, machine learning (ML) algorithms have been considered to learn and find useful and valuable information from large value of data. However, ML algorithms such as neural networks are computationally expensive, and typically, the central processing unit (CPU) is unable to cope with these requirements. Thus, we need a high-performance computer to execute faster solutions such graphics processing unit (GPU). GPUs provide remarkable performance gains compared to CPUs. The GPU is relatively inexpensive with affordable price, availability, and scalability. Since 2006, NVIDIA provides simplification of the GPU programming model with the Compute Unified Device Architecture (CUDA), which supports for accessible programming interfaces and industry-standard languages, such as C and C++. Since then, general-purpose graphics processing unit (GPGPU) using ML algorithms are applied on various applications, including signal and image pattern classification in biomedical area. The importance of fast analysis of detecting cancer or non-cancer becomes the motivation of this study. Accordingly, we proposed soft computing methods, self-organizing map (SOM) and multiple back-propagation (MBP) for big data, particularly on biomedical classification problems. Big data such as gene expression datasets are executed on high-performance computer and Fermi architecture graphics hardware. Based on the experiment, MBP and SOM with GPU-Tesla generate faster computing times than high-performance computer with feasible results in terms of speed and classification performance.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/nssmic.2011.6152953
Implementing Geant4 on GPU for medical applications
  • Oct 1, 2011
  • Hector Perez-Ponce + 8 more

Monte Carlo simulation (MCS) plays a key role in medical applications, especially for emission tomography (ET) and radiotherapy (RT). Unfortunately MCS is also associated with long calculation times that prevent for using it in routine clinical practice. Actually, a solution based on the use of computer clusters to solve the intensive computational issues is not realistic within routine clinical environment. Recently graphics processing units (GPU) became in many domains a cheap solution for the acquisition of a high power computation. The objective of this work was to develop an efficient framework for the implementation of MCS on GPU architectures. Geant4 was used as the MCS engine for targeting medical imaging and radiotherapy applications. We propose the definition of a global strategy and associated structures for such a GPU based simulation. The different steps needed for a Geant4 simulation were implemented on GPU. The first validations have shown equivalence in the underlying photon physics processes between the Geant4 and the GPU codes. Based on these simplistic simulations, we are expecting a speedup factor of over 200 for a complete simulation in emission tomography or in radiotherapy dosimetry.

  • Research Article
  • Cite Count Icon 2
  • 10.1118/1.4813948
SU‐C‐500‐03: GGEMS‐Brachy: Fully GPU Geant4‐Based Efficient Monte Carlo Simulation for Brachytherapy Applications
  • Jun 1, 2013
  • Medical Physics
  • Y Lemarechal + 4 more

Purpose: In brachytherapy, dosimetric plans are routinely calculated with the TG43 formalism which considers the patient as a simple water box. However, an accurate modeling of the physical processes considering patient heterogeneity using Monte Carlo (MC) methods is currently too time‐consuming and computationally demanding to be routinely used. As solution we implemented an accurate and fast MC simulation on graphics processing unit (GPU) for brachytherapy (HDR and LDR) applications. Methods: Based on Geant4 a MC simulation framework was developed on GPU. This framework was extended to include a hybrid GPU navigator, allowing navigation within a voxelized phantom derived from CT imaging including an analytical based structure for accurately modeling the 125I seeds (Source Tech Medical STM1251). In addition, dose scoring based on TLE including uncertainty calculations was incorporated. The implemented full GPU based modeling was compared with a classical CPU MC simulation based on GATE/GEANT4, as well as previously proposed GPU approaches based on the use of phasespace files for the seeds and their positioning based on simple replacement of voxels within the CT volumes. Results: Energy distribution from the seed and dose mapping including uncertainty were compared showing a high agreement (differences <1%). Preliminary results have shown that the GPU implementation is faster by at least two orders of magnitude compared to the GATE mono‐CPU version. A comparison between dosimetric plans based on TG43 and a full MC simulation using the GPU code for LDR prostate brachytherapy led to 40%–100% differences depending on the level of tissue heterogeneity. Conclusion: We propose a full GPU MC simulation based on Geant4, with a hybrid navigator dedicated for brachytherapy applications. Our evaluation shows large dose differences compared to the simplistic TG43 formalism in LDR and very fast execution times compatible with clinical practice.

  • Research Article
  • Cite Count Icon 35
  • 10.1016/j.jcp.2013.04.030
GPU-accelerated Monte Carlo simulation of particle coagulation based on the inverse method
  • May 3, 2013
  • Journal of Computational Physics
  • J Wei + 1 more

GPU-accelerated Monte Carlo simulation of particle coagulation based on the inverse method

  • Conference Article
  • 10.51843/wsproceedings.2018.14
Speeding up Monte Carlo Computations by Parallel Processing Using GPU for Uncertainty Evaluation in Accordance with GUM Supplement 2
  • Jan 1, 2018
  • C.M Tsui + 2 more

The GUM Supplement 2 deals with measurement models with more than one output quantities, which may be mutually correlated. Such measurement models are common in electrical metrology where the measurand can be complex-valued quantities, such as S-parameters. The GUM Supplement 2 describes a Monte Carlo Method (MCM) for evaluating the output quantities, their standard uncertainties, the covariances between them and the coverage region. The Standards and Calibration Laboratory (SCL) has developed six years ago a software tool for evaluation of measurement models for complex-valued quantities in accordance with GUM Supplement 2. The SCL software tool was written in Visual C++ and Visual Basic for Application (VBA), with Microsoft Excel as frontend user interface. As MCM involves large number of repetitive computations, this old SCL software tool has long processing time especially for complicated measurement models such as coaxial airline. Nowadays many personal computers are equipped with Graphics Processing Unit (GPU) containing massive number of floating point cores. A high end GPU may have nearly 2000 cores while the main CPU normally has only up to 4 cores. As MCM is well suited to parallel processing, to speed up the uncertainty computation, SCL has ported the algorithm to GPU using the Open Computing Language (OpenCL) which was specially designed to support parallel computing. The new SCL tool is an add-on module to Microsoft Excel which allows uncertainty budget listed in spreadsheet table to be calculated by MCM. GPU from the major suppliers Nvidia, AMD and Intel are supported. The uncertainty computation time can be reduced by more than ten times. This paper describes the design and implementation of this new software tool.

  • Conference Article
  • Cite Count Icon 7
  • 10.1109/pdgc.2012.6449943
GPU accelerated Monte Carlo simulation of deep penetration neutron transport
  • Dec 1, 2012
  • Bo Yang + 4 more

Over the last decade, with the increasing performance and programmability of Graphics processing unit (GPU), these units have evolved from specialty hardware to massively parallel general computation devices. Simulation of neutron transport plays an important role in national economical construction and large-scale computing in science and engineering. MC (Monte Carlo) simulation of neutron transport owns great advantage over the determined methods to solve some complex types of particle transport. It is the disadvantage that the computational complexity of MC method is very huge. Due to the independence of samples in MC simulation, the algorithm of MC simulation is in principle well-suited to run on highly parallel GPU. However, the complexities of MC simulation of deep penetration particle transport bring serious difficulties in designing a GPU-based algorithm. We present an algorithm based GPU for MC deep penetration particle transport, in which a particle number based task decomposition method and high efficiency parallel data structure are proposed to match with the underlying GPU architecture. Results demonstrate that with the same computational accuracy as MCNP, MCNP-GPU referred to as MCNP integrated with our algorithm on M2050 achieves 3.53-fold and 7.26-fold speedup respectively by compared with MCNP running on X5670 and X5355.

  • Research Article
  • Cite Count Icon 3
  • 10.1115/1.4052078
High-Performance Computing Probabilistic Fracture Mechanics Implementation for Gas Turbine Rotor Disks on Distributed Architectures Including Graphics Processing Units
  • Oct 13, 2021
  • Journal of Engineering for Gas Turbines and Power
  • Mrugesh Gajjar1 + 2 more

We present an efficient Monte Carlo-based probabilistic fracture mechanics simulation implementation for heterogeneous high-performance computing (HPC) architectures including central processing units (CPUs) and graphics processing units (GPUs). The specific application focuses on large heavy-duty gas turbine rotor components for the energy sector. A reliable probabilistic risk quantification requires the simulation of millions to billions of Monte Carlo (MC) samples. We apply a modified Runge–Kutta algorithm in order to solve numerically the fatigue crack growth for this large number of cracks for varying initial crack sizes, locations, material, and service conditions. This compute intensive simulation has already been demonstrated to perform efficiently and scalable on parallel and distributed HPC architectures including hundreds of CPUs utilizing the message passing interface (MPI) paradigm. In this work, we go a step further and include GPUs in our parallelization strategy. We develop a load distribution scheme to share one or more GPUs on compute nodes distributed over a network. We detail technical challenges and solution strategies in performing the simulations on GPUs efficiently. We show that the key computation of the modified Runge–Kutta integration step speeds up over two orders of magnitude on a typical GPU compared to a single threaded CPU. This is supported by our use of GPU textures for efficient interpolation of multidimensional tables utilized in the implementation. We demonstrate weak and strong scaling of our GPU implementation, i.e., that we can efficiently utilize a large number of GPUs/CPUs in order to solve for more MC samples, or reduce the computational turn-around time, respectively. On seven different GPUs spanning four generations, the presented probabilistic fracture mechanics simulation tool ProbFM achieves a speed-up ranging from 16.4× to 47.4× compared to single threaded CPU implementation.

  • Research Article
  • Cite Count Icon 9
  • 10.1016/j.jqsrt.2021.107680
A fast GPU Monte Carlo implementation for radiative heat transfer in graded-index media
  • Apr 27, 2021
  • Journal of Quantitative Spectroscopy and Radiative Transfer
  • Jiang Shao + 2 more

A fast GPU Monte Carlo implementation for radiative heat transfer in graded-index media

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/geoinformatics.2010.5567882
Accelerating spatial clustering detection of epidemic disease with graphics processing unit
  • Jun 1, 2010
  • Sisi Zhao + 1 more

The statistics of disease clustering is of interest to epidemiologists. In order to detect spatial clustering of disease in all the regions of China, we adopted a likelihood ratio based method which utilizes Monte Carlo simulation and spatial exploring to analyze the real time updating data stored in database. However, large number of random tests for Monte Carlo simulation and large scale of the data set had made the speed of analysis too slow to detect and monitor potential public health hazards. Therefore, we explored to adopt graphics processing unit (GPU) and compute unified device architecture (CUDA) to accelerate the spatial exploring and analyzing process. The algorithm has been implemented efficiently on GPU and the access pattern to memory has been optimized to exploit the computing power of GPU. As a result, the GPU based spatial exploring and likelihood ratio test program performed more than forty times faster then the CPU implementation. The Monte Carlo simulation on GPU performed around thirty times faster than the counter part on CPU. By using GPU and CUDA, the usage of our application is changed from verification after the event to early warning.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant