Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Estimasi Cadangan Klaim Individu untuk Klaim Short-Tailed dan Long-Tailed Menggunakan Algoritma Backpropagation

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

In insurance, risk can occur at any time, causing claims to sometimes have a large amount of value, so that insurance companies may not be able to satisfy claim payments. If these situations occur, insurance companies need claim reserves to prepare for such events. There are several methods to calculate claim reserves, such as aggregate claim reserving. However, certain claim characteristics involve dependencies among claims, which result in a lack of detailed information for individual claims. In addition, an increasing number of claims becomes more difficult to compute using traditional methods. Therefore, this research aims to calculate individual claim reserves using one branch of machine learning, namely the Backpropagation Algorithm. The Backpropagation Algorithm is believed to remain relevant compared to other algorithmic models because, in several studies, it produces relatively low values of Mean Absolute Percentage Error (MAPE), at approximately 2.70%. The data used in this research are simulated using R software, generating 10,000 claims over 20 years, consisting of 6,000 short-tailed claims and 4,000 long-tailed claims. The data model is evaluated using MAPE. The resulting MAPE value is 0.55%, indicating that the data are highly suitable for predictive modeling. The prediction results show that the total claims to be paid in the 21st development year reach Rp22,945,450,000,000, with an average claim amount of approximately Rp2,294,545,152. This research contributes to both informatics and actuarial science by developing an individual claim reserving approach to predict claim payments more efficiently.

Similar Papers
  • Research Article
  • 10.24912/jiksi.v11i2.26002
PREDIKSI CURAH HUJAN DI KABUPATEN BADUNG, BALI MENGGUNAKAN METODE LONG SHORT-TERM MEMORY
  • Aug 23, 2023
  • Jurnal Ilmu Komputer dan Sistem Informasi
  • Brando Dharma Saputra + 2 more

Rainfall is the height of rainwater that falls on a flat area, assuming it doesn't evaporate, doesn't seep, and doesn't flow. Rain levels are measured in mm (millimeters). The target of the research being conducted is in Badung Regency, Bali because Bali is a tourist area that is often visited by tourists and from Indonesian itself, so predictions of meteorology, such as rainfall will greatly impact tourism. In this test, predictions use the Long Short Term Memory (LSTM) method, using daily weather data from the BMKG from 2010 to 2020 as training data and daily weather data for 2021 as prediction data. Based on the test results above, the results show that the two LSTM tests with LSTM Model 128.64 and LSTM Model 64.32 have low MAE and MAPE error values. From First Scenario, the Mean Absolute Error (MAE) value is 8.97246598930908 and Mean Absolute Percentage Error (MAPE) value is 1.7657206683278308%. From Second Scenario, the Mean Absolute Error is 9.706669940783014 and Mean Absolute Percentage Error is 1.9028466692362323%. From the MAE and MAPE values obtained in these two scenarios, it can be proven that from the evaluation results of Rainfall predictions in Badung Regency, Bali, the predictions can be said to be very accurate because they have an error value of less than 10.

  • Research Article
  • Cite Count Icon 1
  • 10.5152/electrica.2024.22107
Short-Term Electricity Consumption Forecasting at Residential Level Using Two-Phase Hybrid Machine Learning Model
  • Apr 17, 2024
  • ELECTRICA
  • Alisha Banga + 1 more

The need for electricity has increased as the population, electrical appliances, electric cars, and industrialization have grown. Therefore, an accurate short-term electricity forecasting is required because it is helpful in day-to-day scheduling activities of utility companies like transmission and generation of electric energy. It can make the power grid safe, reduce electricity production costs, and fulfil user needs and economic and environment benefits. A single model may not be able to solve the electricity consumption problem because it contains linear and non-linear data. In this study, a two-phase hybrid machine learning model is developed for electricity utilisation prediction at a residential level. In the first phase, two algorithms, namely extreme gradient-boosting and linear regression are combined to learn the trend, seasonality, randomness, and cyclic components of the data. In the second phase, a voting ensemble model optimized is applied considering the best three models from 11 baseline models and weight parameter of the models is optimized using genetic algorithm. The developed model outperformed all models (baseline and state-of-the-art) considering four performance parameters over the individual household electric power consumption dataset. The proposed model has given Mean Squared Error (MSE) value of 0.025, Room Mean Squared Error (RMSE) value of 0.162, Mean Absolute Error (MAE) value of 0.129, and Mean Absolute Percentage Error (MAPE) value of 15.61 on a daily-level dataset. The proposed model has given a 0.159 MSE value, 0.387 RMSE value, 0.283 MAE value, and 25.07 MAPE value on an hourly-level dataset. Analysis of variance one-way statistical test is applied to show that results are statistically significant. Cite this article as: A. Banga and S. C. Sharma, “Short-term electricity consumption forecasting at residential level using two-phase hybrid machine learning model,” Electrica, 24(2), 272-283, 2024.

  • Research Article
  • Cite Count Icon 1
  • 10.47065/josyc.v5i4.5810
Perbandingan Metode Double Exponential Smoothing dan Double Moving Average dalam Penjualan Produk Herbal HNI
  • Aug 28, 2024
  • Journal of Computer System and Informatics (JoSYC)
  • Sunilfa Maharani Tanjung + 1 more

Predicting sales of herbal products is the main challenge faced in this research. Currently, forecasting is done solely based on previous sales records, which often prove to be inaccurate. As a result, the company frequently experiences unstable sales and faces difficulties in optimizing inventory and planning product orders efficiently to meet high customer demand. This situation often leads to significant cost losses, forcing the company to reduce capital costs for certain products to cover these losses. The problem arises because the company has not yet implemented an appropriate forecasting method, resulting in estimates that are not supported by a reliable system. This research aims to design and implement a Sales Management Information System for Herbal Products, focusing on the use of the Double Exponential Smoothing (DES) and Double Moving Average (DMA) methods. Additionally, this study aims to compare the two methods in predicting sales by analyzing and calculating the Mean Absolute Percentage Error (MAPE) value for each method. The research uses 200 sales data points from 8 best-selling products, including one of the best-selling products, HNI HEALTH, with data collected from April 2022 to April 2024. The MAPE results from the sales data were then calculated by summing all the MAPE values and dividing them according to the number of MAPEs with an alpha of 0.3. It was found that the DES method is more accurate, with an average Mean Absolute Percentage Error (MAPE) value of 0.285, compared to the DMA method, which has an average MAPE value of 0.292. The DES method is considered more accurate because its MAPE value is smaller than that of the DMA method.

  • Research Article
  • Cite Count Icon 17
  • 10.1111/njb.02768
Artificial neural networking to estimate the leaf area for invasive plant Wedelia trilobata
  • Jun 1, 2020
  • Nordic Journal of Botany
  • Ahmad Azeem + 3 more

Leaf area are very important parameter for the understanding of growth and physiological responses of invasive plant species under different environmental factors. This study was conducted to build non‐destructive leaf area model of Wedelia trilobata that were grown in greenhouse. Regression analysis and artificial neural network (ANN) approaches were used for the development of leaf area model with the help of leaf length and width of 262 plants samples. In selection of best method under both techniques, the lower value of mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE) and higher value of R 2 were considered. According to the results it was found that ANN have higher value of (R 2 = 0.96) and lower value of error (MAE = 0.023, RMSE = 0.379, MAPE = 0.001) than regression analysis (R 2 = 0.94, MAE = 0.111, RMSE = 1.798, MAPE = 0.0005). It was concluded that error between predicted and actual value was less under ANN. Therefore, ANN model approach can be used as an alternating method for the estimation of leaf area. Through estimation of leaf area, invasive plant growth can predict under different environment conditions.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 14
  • 10.1038/s41598-025-91878-0
Seasonal forecasting of the hourly electricity demand applying machine and deep learning algorithms impact analysis of different factors
  • Mar 18, 2025
  • Scientific Reports
  • Heba-Allah Ibrahim El-Azab + 3 more

The purpose of this paper is to suggest short-term Seasonal forecasting for hourly electricity demand in the New England Control Area (ISO-NE-CA). Precision improvements are also considered when creating a model. Where the whole database is split into four seasons based on demand patterns. This article’s integrated model is built on techniques for machine and deep learning methods: Adaptive Neural-based Fuzzy Inference System, Long Short-Term Memory, Gated Recurrent Units, and Artificial Neural Networks. The linear relationship between temperature and electricity consumption makes the relationship noteworthy. Comparing the temperature effect in a working day and a temperature effect on a weekend day where at night, the marginal effects of temperature on the demand in a working day for power are likewise at their highest. However, there are significant effects of temperature on the demand for a holiday, even a weekend or special holiday. Two scenarios are used to get the results by using machine and deep learning techniques in four seasons. The first scenario is to forecast a working day, and the second scenario is to forecast a holiday (weekend or special holiday) under the effect of the temperature in each of the four seasons and the cost of electricity. To clarify the four techniques’ performance and effectiveness, the results were compared using the Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), Normalized Root Mean Squared Error (NRMSE), and Mean Absolute Percentage Error (MAPE) values. The forecasting model shows that the four highlighted algorithms perform well with minimal inaccuracy. Where the highest and the lowest accuracy for the first scenario are (99.90%) in the winter by simulating an Adaptive Neural-based Fuzzy Inference System and (70.20%) in the autumn by simulating Artificial Neural Network. For the second scenario, the highest and the lowest accuracy are (96.50%) in the autumn by simulating Adaptive Neural-based Fuzzy Inference System and (68.40%) in the spring by simulating Long Short-Term Memory. In addition, the highest and the lowest values of Mean Absolute Error (MAE) for the first scenario are (46.6514, and 24.759 MWh) in the spring, and the summer by simulating Artificial Neural Networks. The highest and the lowest values of Mean Absolute Error (MAE) for the second scenario are (190.880, and 45.945 MWh) in the winter, and the autumn by simulating Long Short-Term Memory, and Adaptive Neural-based Fuzzy Inference System.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 21
  • 10.3390/aerospace9070394
Flight Departure Time Prediction Based on Deep Learning
  • Jul 21, 2022
  • Aerospace
  • Hang Zhou + 4 more

Accurate flight departure time prediction enables the rational use of airport support resources, aprons, and runway resources, and promotes the implementation of collaborative decision-making. In order to accurately predict the flight departure time, this paper proposes a deep learning-based flight departure time prediction model. First, this paper analyzes the influence of different factors on flight departure time and the influencing factor. Secondly, this paper establishes a gated recurrent unit (GRU) model, considers the impact of different hyperparameters on network performance, and determines the optimal hyperparameter combination through parameter tuning. Finally, the model verification and comparative analysis are carried out using the real flight data of ZSNJ. The evaluation values of the established model are as follows: root mean square error (RMSE) value is 0.42, mean absolute percentage error (MAPE) value is 6.07, and mean absolute error (MAE) value is 0.3. Compared with other delay prediction models, the model established in this paper has a 16% reduction in RMSE, 34% reduction in MAPE, and 86% reduction in MAE. The model has high prediction accuracy, which can provide a reliable basis for the implementation of airport scheduling and collaborative decision-making.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.29207/resti.v6i4.4134
Support Vector Regression Method for Predicting Off-Grid Photovoltaic Output Power in the Short Term
  • Aug 22, 2022
  • Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi)
  • Kharisma Bani Adam + 3 more

Photovoltaic (PV) technology is a renewable technology utilizing conversion of solar power or solar radiation into electrical energy. In the manufacture of Solar Power Generation systems, reference is needed regarding the cost of generation and scheduling of maintenance plans. To obtain this reference, it is necessary to predict the photovoltaic power output which is used to determine the power output of PV in the future. In this study, a system that is used to predict short-term power output in PV is designed. This system uses solar irradiation data and 42 days of power output in off-grid PV mini-grid as the dataset. The dataset obtained from the PV output is processed using the Support Vector Regression method with the Kernel Radial Basis Function (RBF) function. Based on the dataset used, this study succeeded in testing the best kernel, namely the RBF kernel. Evaluation of the prediction model obtained a smaller error value than other kernel tests with a Mean Absolute Percentage Error (MAPE) value of 21.082%, Mean Square Error (MSE) value of 0.122, and Mean Absolute Error (MAE) value of 0.262. The prediction model obtained is used to predict the short-term PV power output for the next 3 days. The results of the prediction model have an error value of 5.785 % for MAPE, 0.005 for MAE and 0.069 for MSE. Therefore, the predictive model can be categorized as very good and feasible to predict short-term power output

  • Research Article
  • Cite Count Icon 100
  • 10.1016/j.fuel.2016.03.028
Experimental and regression analysis of noise and vibration of a compression ignition engine fuelled with various biodiesels
  • Mar 16, 2016
  • Fuel
  • Erinç Uludamar + 2 more

Experimental and regression analysis of noise and vibration of a compression ignition engine fuelled with various biodiesels

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 12
  • 10.4067/s0718-221x2018005041501
A modeling study to evaluate the quality of wood surface
  • Jan 1, 2018
  • Maderas. Ciencia y tecnología
  • Ender Hazir + 1 more

The goal of this study was to develop a model to predict sanding conditions of different type of materials such as Lebnon cedar (Cedrus libani) and European Black pine (Pinus nigra). Specimens were prepared using different values of grit size, cutting speed, feed rate, and sanding direction. Surface quality values of specimens were measured employing a laser- based robotic measurement system and stylus type measurement equipment. Full factorial design based Analysis of Variance was applied to determine the effective factors. These factors were used to develop the Artificial Neural Networks models for two different measurement systems. The MATLAB Neural Network Toolbox was used to predict the Artificial Neural Networks models. According to the results, the Artificial Neural Networks models were performed using Mean Absolute Percentage Error and R-square values. Mean Absolute Percentage Error values for laser and stylus equipment were found as 2.405 % and 3.766 %, respectively. R-square values were determined as 96.2% and 92.7 % for laser and stylus measurement equipment, respectively. These results showed that the proposed models can be successfully used to predict the surface roughness values.

  • Research Article
  • Cite Count Icon 12
  • 10.1080/01932691.2024.2339458
Adsorptive uptake of thymol blue from aqueous medium using calcined snail shells: equilibrium, kinetic, thermodynamic, neuro-fuzzy and DFT studies
  • Apr 4, 2024
  • Journal of Dispersion Science and Technology
  • Victoria Faith Gold + 9 more

ABSTRACTS The spontaneous discharge of pollutants or toxicants into the ecosystem as a result of various anthropogenic activities is alarming, this necessitates a drastic cheap, and eco-friendly clean-up approach. Out of all the methods, adsorption has proven to be the most effective. The batch adsorptive removal of Thymol Blue (TB) dye from aqueous solutions is examined in this study utilizing calcined snail shells (CalSS). With the aid of SEM, EDS, XRD, and FTIR, the synthesized adsorbent was examined for different physicochemical characteristics. A machine learning model namely Adaptive Neuro-Fuzzy Inference System (ANFIS) was developed to assess the TB adsorption process while taking into account the important adsorption parameters contact time, temperature, pH, adsorbent dosage, then initial adsorbate concentration. Adsorbent pHpzc was also evaluated. EDX and FTIR confirm the formation of CaO with sharp peaks at 547 cm−1 and C–O and O–H are present. SEM and XRD show an irregularly shaped highly crystalline adsorbent material having a typical particle size of 65 ± 2.81 nm and lattice parameter value of 8.611617 Å. The pHpzc value is 11.04, indicating basic surface characteristics. The pH of 3.0, an adsorbent dose of 10 mg and the highest achievable adsorption efficiency were measured to be 98.75% at 20 °C. The findings from the study fit nicely onto Brouser Sotolongo with q BS = 259.4887 mg/g and R 2 = 0.9888. The pseudosecond order model (PSOM) recorded the least error value of 0.3336 and R 2 = 0.9952. This indicates chemisorption and multilayer adsorption processes. The thermodynamic parameters ΔH° and ΔG° demonstrate the exothermic and spontaneous nature of the adsorption process. The ANFIS model for the TB sequestration was evaluated using relevant statistical metrics, giving Root Mean Square Error (RMSE) value of 1.4644, Mean Absolute Deviation (MAD) value of 0.576, Mean Absolute Error (MAE) value of 1.2974 and Mean Absolute Percentage Error (MAPE) value of 1.5141. This outcome revealed that the ANFIS model and experimental findings are in good agreement. The efficacy of calcined snail shell in removing Thymol Blue dye from aqueous solution was also examined using in-silico technique which was in accordance to the value of the calculated descriptors obtained from the adsorbent as well as the calculated binding affinity. The elimination of TB-polluted wastewater using calcined snail shells is demonstrated in this study to be a successful and environmentally benign process.

  • Research Article
  • Cite Count Icon 22
  • 10.1039/d4ra01074d
Thermally modified nanocrystalline snail shell adsorbent for methylene blue sequestration: equilibrium, kinetic, thermodynamic, artificial intelligence, and DFT studies
  • Jan 1, 2024
  • RSC Advances
  • Abisoye Abidemi Adaramaja + 9 more

In recent years, the quest for an efficient and sustainable adsorbent material that can effectively remove harmful and hazardous dyes from industrial effluent has become more intense. The goal is to explore the capability of thermally modified nanocrystalline snail shells (TMNSS) as a new biosorbent for removing methylene blue (MB) dye from contaminated wastewater. TMNSS was employed in batch adsorption experiments to remove MB dye from its solutions, taking into account various adsorption parameters such as contact time, temperature, pH, adsorbent dosage, and initial concentration. SEM, EDS, XRD, and FTIR were used to characterize the adsorbent. The study further developed and adopted adaptive neuro-fuzzy inference system (ANFIS) and density functional theory (DFT) studies to holistically examine the adsorption process of MB onto the adsorbent. EDX and FTIR confirm the formation of CaO with a sharp peak at 547 cm−1, and C–O and O–H are present, as well. SEM and XRD show an irregularly shaped highly crystalline nanosized (65 ± 2.81 nm) particle with a lattice parameter value of 8.611617 Å. The adsorption efficiency of 96.48 ± 0.58% was recorded with a pH of 3.0 and an adsorbent dose of 10 mg at 30 °C. The findings from the study fit nicely onto Freundlich isotherms, with Qm = 31.7853 mg g−1 and R2 = 0.9985. Pseudo-second-order kinetics recorded the least error value of 0.8792 and R2 = 0.9868, thus indicating chemisorption and multilayer adsorption processes. The exothermic and spontaneous nature of the adsorption process are demonstrated by ΔH° and ΔG°. The performance of the ANFIS-based prediction of removal rate, which was demonstrated by a root mean square error (RMSE) value of 2.2077, mean absolute deviation (MAD) value of 1.1429, mean absolute error (MAE) value of 1.8786, and mean absolute percentage error (MAPE) value of 2.0178, revealed that the ANFIS model predictions and experimental findings are in good agreement. More so, DFT provides insights into the molecular interactions between MB and the adsorbent surface, with a calculated adsorbate–adsorbent binding affinity value of −1.3 kcal mol−1, thus confirming the ability of TMNSS for MB sequestration. The findings of this study highlight the promising potential of thermally modified nanocrystalline snail shells as sustainable and efficient adsorbents for MB sequestration.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 12
  • 10.4209/aaqr.230006
Forecasting PM2.5 in Malaysia Using a Hybrid Model
  • Jun 5, 2023
  • Aerosol and Air Quality Research
  • Ezahtulsyahreen Ab Rahman + 3 more

Predicting future PM2.5 concentrations based on knowledge obtained from past observational data is very useful for predicting air pollution. This paper aims to develop a hybrid forecasting model using an Artificial Neural Network (ANN) and Triple Exponential Smoothing (TES) on clustered PM2.5 data from a HPR (High Pollution Region), MPR (Medium Pollution Region), and LPR (Low Pollution Region) in Malaysia. Historical PM2.5 concentrations in Malaysia from January 2018 to December 2019 were used to develop a hybrid model. The proposed hybrid model was then evaluated in terms of Mean Absolute Percentage Error (MAPE) values by comparing them with real PM2.5 data from the year 2020 in the HPR, MPR and LPR. The results showed that the hybrid model of ANN and TES presented the lowest RMSE (Root Mean Squared Error) (4.25–8.56 µg m−3), MAE (Mean Absolute Error) (2.51–4.95 µg m−3), MAPE (0.13–0.2%), and MASE (Mean Absolute Scaled Error) (1.45–2.01) in different areas of pollution compared with other models. The comparison between the ANN and TES hybrid models and the real PM2.5 data in 2020 showed that the models gave sufficient accuracy in the HPR and MPR with MAPE values of between 20% and 50%, while the LPR showed less accuracy due to the high value of MAPE of more than 50%. Overall, the hybrid model developed in this study opens up a new prediction method for air quality forecasting and is sufficiently accurate to be used as a tool for air quality management.

  • Research Article
  • Cite Count Icon 21
  • 10.1007/s40710-020-00431-w
Artificial Neural Network (ANN) Modelling of Palm Oil Mill Effluent (POME) Treatment with Natural Bio-coagulants
  • Apr 24, 2020
  • Environmental Processes
  • Nurul Asyikin Mohd Najib + 3 more

Raw palm oil mill effluent (POME) is classified as a highly polluting effluent, which needs to be treated to an acceptable level before being discharged into water bodies. Currently, chemical coagulants are widely used in treating POME, but their hazardous nature has caused several health and environmental problems. Therefore, this research presents the use of natural materials such fenugreek and okra as bio-coagulants and bio-flocculants, respectively, for the treatment of POME. Artificial neural network (ANN) modelling technique was used for the estimation of predicted results of the coagulation-flocculation process. The responses of the process were the percentage removal of total suspended solid (TSS), turbidity (TUR) and chemical oxygen demand (COD), while the inputs were fenugreek dosages, okra dosages, pH and mixing speed. The ANN model was developed using 12 different training algorithms. Scaled conjugate algorithm (SCG) proved to be the best training algorithm with lowest mean-squared error (MSE), mean absolute error (MAE), mean absolute percentage error (MAPE) values for all the three outputs. The MSE, MAE and MAPE values were: 16.64, 2.10, 0.03 for the TSS response; 5.05, 1.03, 0.01 for the TUR response; and 54.59, 3.82, 0.07 for the COD response, respectively. The ANN model with a regression R of 0.8629 proved to be the best for the prediction of all the responses in this study. The results proved that ANN can be applied to predict TSS, TUR and COD of POME.

  • Research Article
  • 10.1038/s41598-026-35754-5
The peak shifting electricity consumption management and influencing factors of smart grid from recurrent neural network model and deep learning.
  • Jan 16, 2026
  • Scientific reports
  • Feng Wang + 2 more

This study uses Recurrent Neural Network (RNN) model and Electricity Consumption (EC) management model to conduct in-depth research and analysis to achieve effective management of peak shifting EC in smart grids and analyze its influencing factors. A management method that is superior to other models has been found, which has a breakthrough effect on the use and research of smart grids. Firstly, the RNN model is analyzed, and a smart grid energy consumption prediction model based on EC management of peak shifting is established. Secondly, with the goal of shifting peaks and filling valleys, the power load data with several different characteristics is divided into different EC modes. Finally, in predicting smart grid energy consumption, it is compared with Linear Regression (LR), nonlinear regression, and Autoregressive Integrated Moving Average Mode prediction models. The results show that: (1) Before using the hydrogen energy peak shifting and valley filling EC management mode, the peak energy consumption is about 46kWh, and the adjusted peak energy consumption decreases by about 12kWh. This indicates that the hydrogen peak shifting EC management mode operates well in smart grids. (2) In practical environments, the Mean Square Error (MSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE) values of the RNN-based energy consumption time series prediction model are 0.124, 0.26, and 5.75, respectively. The MSE, MAE, and MAPE values of the LR prediction model are 22.09, 3.33, and 21.48, respectively. The MSE, MAE, and MAPE values of the nonlinear regression prediction model are 0.223, 489, and 16.32, respectively, indicating that the prediction accuracy of the RNN energy consumption time series prediction model is superior to other comparative models. This study has important reference value for schools and factories to use artificial intelligence to elucidate future energy consumption needs and reasonable allocation of energy in the power grid.

  • Research Article
  • Cite Count Icon 23
  • 10.1038/s41598-025-95709-0
Well log data generation and imputation using sequence based generative adversarial networks
  • Mar 31, 2025
  • Scientific Reports
  • Abdulrahman Al-Fakih + 4 more

Well log analysis is significant for hydrocarbon exploration, providing detailed insights into subsurface geological formations. However, gaps and inaccuracies in well log data, often due to equipment limitations, operational challenges, and harsh subsurface conditions, can introduce significant uncertainties in reservoir evaluation. Addressing these challenges requires effective methods for both synthetic data generation and precise imputation of missing data, ensuring data completeness and reliability. This study introduces a novel framework utilizing sequence-based generative adversarial networks (GANs) specifically designed for well log data generation and imputation. The framework integrates two distinct sequence-based GAN models: time series GAN (TSGAN) for generating synthetic well log data and sequence GAN (SeqGAN) for imputing missing data. Both models were tested on a dataset from the North Sea, Netherlands region. For the imputation task, the input comprises logs with missing values and the output is the corresponding imputed logs; for the synthetic data generation task, the input is complete real logs and the output is synthetic logs that mimic the statistical properties of the original data. All log measurements are normalized to a 0-1 range using min-max scaling, and error metrics are reported in these normalized units. Different sections of 5, 10, and 50 data points were used. Experimental results demonstrate that this approach achieves superior accuracy in filling data gaps compared to other deep learning models for spatial series analysis. The imputation method yielded values of 0.92, 0.86, and 0.57, with corresponding mean absolute percentage error (MAPE) values of 8.320, 0.005, and 166.6, and mean absolute error (MAE) values of 0.012, 0.002, and 0.03, respectively. The synthetic generation yielded of 0.92, MAE, of 0.35, and MRLE of 0.01. These results set a new benchmark for data integrity and utility in geosciences, particularly in well log data analysis.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant