Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Formation of machine learning models of high-refractive index organic materials through high-throughput quantum chemical calculations

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

We studied the machine learning models for high-refractive organic compounds, based on literature data. Although, the prediction data by the model were well toward common organic compounds, that of the high-refractive compounds were not accord with experimental data, due to the few literature data of high-refractive compounds. Therefore, we created data for 62 high-refractive compounds using high-throughput quantum chemical calculations and performed the transfer learning based on the literature data model. The obtained transfer learning model showed well results for selected high-refractive organic compounds.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 4
  • 10.1038/s41598-022-20012-1
Perception without preconception: comparison between the human and machine learner in recognition of tissues from histological sections
  • Sep 30, 2022
  • Scientific Reports
  • Sanghita Barui + 4 more

Deep neural networks (DNNs) have shown success in image classification, with high accuracy in recognition of everyday objects. Performance of DNNs has traditionally been measured assuming human accuracy is perfect. In specific problem domains, however, human accuracy is less than perfect and a comparison between humans and machine learning (ML) models can be performed. In recognising everyday objects, humans have the advantage of a lifetime of experience, whereas DNN models are trained only with a limited image dataset. We have tried to compare performance of human learners and two DNN models on an image dataset which is novel to both, i.e. histological images. We thus aim to eliminate the advantage of prior experience that humans have over DNN models in image classification. Ten classes of tissues were randomly selected from the undergraduate first year histology curriculum of a Medical School in North India. Two machine learning (ML) models were developed based on the VGG16 (VML) and Inception V2 (IML) DNNs, using transfer learning, to produce a 10-class classifier. One thousand (1000) images belonging to the ten classes (i.e. 100 images from each class) were split into training (700) and validation (300) sets. After training, the VML and IML model achieved 85.67 and 89% accuracy on the validation set, respectively. The training set was also circulated to medical students (MS) of the college for a week. An online quiz, consisting of a random selection of 100 images from the validation set, was conducted on students (after obtaining informed consent) who volunteered for the study. 66 students participated in the quiz, providing 6557 responses. In addition, we prepared a set of 10 images which belonged to different classes of tissue, not present in training set (i.e. out of training scope or OTS images). A second quiz was conducted on medical students with OTS images, and the ML models were also run on these OTS images. The overall accuracy of MS in the first quiz was 55.14%. The two ML models were also run on the first quiz questionnaire, producing accuracy between 91 and 93%. The ML models scored more than 80% of medical students. Analysis of confusion matrices of both ML models and all medical students showed dissimilar error profiles. However, when comparing the subset of students who achieved similar accuracy as the ML models, the error profile was also similar. Recognition of ‘stomach’ proved difficult for both humans and ML models. In 04 images in the first quiz set, both VML model and medical students produced highly equivocal responses. Within these images, a pattern of bias was uncovered–the tendency of medical students to misclassify ‘liver’ tissue. The ‘stomach’ class proved most difficult for both MS and VML, producing 34.84% of all errors of MS, and 41.17% of all errors of VML model; however, the IML model committed most errors in recognising the ‘skin’ class (27.5% of all errors). Analysis of the convolution layers of the DNN outlined features in the original image which might have led to misclassification by the VML model. In OTS images, however, the medical students produced better overall score than both ML models, i.e. they successfully recognised patterns of similarity between tissues and could generalise their training to a novel dataset. Our findings suggest that within the scope of training, ML models perform better than 80% medical students with a distinct error profile. However, students who have reached accuracy close to the ML models, tend to replicate the error profile as that of the ML models. This suggests a degree of similarity between how machines and humans extract features from an image. If asked to recognise images outside the scope of training, humans perform better at recognising patterns and likeness between tissues. This suggests that ‘training’ is not the same as ‘learning’, and humans can extend their pattern-based learning to different domains outside of the training set.

  • Research Article
  • Cite Count Icon 33
  • 10.1016/j.eswa.2023.121300
Developing deep transfer and machine learning models of chest X-ray for diagnosing COVID-19 cases using probabilistic single-valued neutrosophic hesitant fuzzy
  • Sep 1, 2023
  • Expert Systems with Applications
  • Hassan A Alsattar + 6 more

Developing deep transfer and machine learning models of chest X-ray for diagnosing COVID-19 cases using probabilistic single-valued neutrosophic hesitant fuzzy

  • Research Article
  • 10.7759/cureus.109468
Comparison of Prognostic Performance Between a Machine Learning Model and Manually Measured Grey-White-Matter Ratio on Early Brain Computed Tomography After Out-of-Hospital Cardiac Arrest
  • May 1, 2026
  • Cureus
  • Fumiya Inoue + 3 more

ObjectivesEarly prediction of neurological outcomes in patients with out-of-hospital cardiac arrest (OHCA) is critical for guiding treatment decisions. Machine learning (ML) model and grey-white matter ratio (GWR), both derived from brain computed tomography (CT), can be used to predict the neurological outcome. However, their relative performance shortly post-return of spontaneous circulation (ROSC) and whether combining the ML model with prehospital information can improve predictive performance remains unclear. This study aimed (1) to compare the predictive performance for poor neurological outcome between the ML model and manually measured GWR in the early phase after ROSC in patients with OHCA, and (2) to assess the predictive ability of the combination of the ML model and prehospital information.MethodsThis single-center retrospective study included adult patients who underwent brain CT within two hours post-ROSC. The endpoint was consecutive coma post-ROSC. Three slice levels (basal ganglia, centrum semiovale, high convexity) of brain CT images were used to generate the ML model and the GWR. Residual Network 101 (ResNet-101) with transfer learning was constructed in the ML model.ResultsAmong the 143 cases, 88 patients had a persistent coma, and 55 awoke from coma. Across 10 repeated five-fold cross-validations, the area under the receiver operating characteristic curves (AUCs) for predicting persistent coma between the three-slice ensembled ML model and the average GWR were not significantly different (ML model: 0.796 (interquartile range (IQR): 0.737-0.826), GWR: 0.821 (IQR: 0.763-0.854); p = 0.121). In the regression analysis, the AUC of the model based on prehospital information was 0.846 (95% CI: 0.772-0.92), which improved to 0.905 (95% CI: 0.856-0.953) after adding the ML score.ConclusionThe ML model achieved moderate predictive performance, with no significant difference compared with the conventional GWR method. The combination of the ML model and prehospital information could improve predictive performance.

  • Research Article
  • 10.1016/j.msea.2025.149714
Assessing machine learning predictions of austenitic steel compositions for optimum mechanical response
  • Feb 1, 2026
  • Materials Science and Engineering: A
  • Linh Thi Hoai Nguyen + 9 more

Understanding and predicting mechanical properties such as proof stress (yield strength), ultimate tensile strength, elongation to failure, and reduction in area are essential for screening the application of austenitic stainless steels in adverse chemomechanical environments. However, experimental determination of these properties is time-consuming, labor-intensive, and costly—especially under extreme conditions, which requires advanced experimental capabilities. In this study, we leverage a large-scale, curated dataset comprising 2180 experimental entries of austenitic alloys to develop machine learning (ML) models capable of predicting these key mechanical properties as functions of composition, solution treatment condition, and testing temperature . We systematically evaluate a range of ML algorithms, namely linear regression, kernel ridge regression, extreme gradient boosting, and artificial neural network. Among these, the extreme gradient boosting achieves the highest predictive accuracy, with R 2 scores of 0.946 and 0.985 for proof stress and ultimate tensile strength, respectively. To further enhance model performance, we explore ensemble learning and transfer learning strategies. The transfer learning approach that leverages interdependencies between mechanical properties reduces the mean percentage error by 21.0% and increases R 2 score by 3.1% in predicting reduction in area, compared to the original artificial neural network model. Our results show that ML models trained on well-structured experimental data can serve as an efficient and reliable tool for exploring the effect of composition on screening metrics. This work highlights the potential of data-driven approaches to accelerate the design and optimization of high-performance austenitic alloys. Furthermore, obtaining feature importance and insights from extreme gradient boosting using Shapley Additive Explanation tool provides valuable understanding of how features contribute to the prediction of mechanical properties. • Developed machine learning models to predict key mechanical properties of austenitic stainless steels as functions of compositions, solution treatment conditions, and testing temperature, leveraging 2180 curated experimental data points. • Achieved high predictive accuracy with R 2 scores of 0.95, 0.99, 0.90, and 0.88 for yield strength, ultimate tensile strength, elongation to failure, and reduction in area, respectively. • Implemented advanced ensemble and transfer learning techniques to further enhance predictive performance beyond conventional ML models. Transfer learning leveraging knowledge from ultimate tensile strength reduced the mean absolute percentage error in predicting reduction in area by 21% compared to the baseline Artificial Neural Network (ANN) model. • Demonstrated the capability of transfer learning to effectively transfer knowledge across related mechanical property domains. • Employed SHapley Additive exPlanations (SHAP) to interpret feature importance, identifying the roles of Nb, Ti, Mo, C, N, and B in influencing strength and ductility, consistent with metallurgical understanding. • Highlighted that the developed ML models successfully capture complex relationships between composition, processing, and mechanical properties, offering a pathway to accelerate austenitic alloy design.

  • Research Article
  • Cite Count Icon 49
  • 10.1016/j.resourpol.2023.104216
A novel deep-learning technique for forecasting oil price volatility using historical prices of five precious metals in context of green financing – A comparison of deep learning, machine learning, and statistical models
  • Oct 1, 2023
  • Resources Policy
  • Muhammad Mohsin + 1 more

A novel deep-learning technique for forecasting oil price volatility using historical prices of five precious metals in context of green financing – A comparison of deep learning, machine learning, and statistical models

  • Research Article
  • Cite Count Icon 42
  • 10.1016/j.cemconcomp.2024.105488
Transfer learning enables prediction of steel corrosion in concrete under natural environments
  • Feb 24, 2024
  • Cement and Concrete Composites
  • Haodong Ji + 3 more

Transfer learning enables prediction of steel corrosion in concrete under natural environments

  • Research Article
  • Cite Count Icon 2
  • 10.17762/turcomat.v12i8.3942
Performance Enhancement of Hybrid Algorithm for Bank Telemarketing
  • Apr 20, 2021
  • Turkish Journal of Computer and Mathematics Education (TURCOMAT)
  • Rohan Desai

Telemarketing is an interactive direct marketing system in which telemarketers encourage customers to leverage the resources by notifying, imparting knowledge of online products, latest business offers via direct interaction or through a telephone call. In the contemporary global pandemic spell telemarketing has become dominant backbone to increase the online banking business to withstand for the reducing retail business. It has gained prominance in the banking and financial sector with the enormous adoption and availability of cellular connections amongst customers. The contemporary work has scrutinized conventional classification as well as data mining methods have a problem of ill-fitting with multiple features and are prone to data leakage during re-training of the machine learning model. A local Indian bank were designated, contemplating the current economic slowdown and crisis. A discussion on three machine learning (ML) models is performed along with the Hybrid ML model, Logistic Regression ML model (LR), Naive Bayes ML model (NB), Decision Trees ML model (DTs). The three ML models were tested and analysed with proposed Hybrid ML model on an evaluation set, the data is partitioned as training, validation and test set. The hybrid model first identifies important features of subscribed customers and predicts response for a potential customer, both existing and new who will eventually subscribe again through the direct marketing campaign. The hybrid model is trained to predict the response of new customer who will subscribe to the product or service offered via a direct marketing campaign through transfer learning. The hybrid model API shows new customer response on the front-end screen. To overcome the problem of ill-fitting and data leakage, the model is trained on a large dataset and tuned on a validation set. The proposed hybrid machine learning technique presented the best results (Accuracy 98.69%). Python language is used to develop the model. Financial institutions and organizations can use the hybrid model for predictions of product direct marketing response with customer transaction information.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 42
  • 10.1007/s11356-024-35764-8
An examination of daily CO2 emissions prediction through a comparative analysis of machine learning, deep learning, and statistical models
  • Jan 1, 2025
  • Environmental Science and Pollution Research
  • Adewole Adetoro Ajala + 3 more

Human-induced global warming, primarily attributed to the rise in atmospheric CO2, poses a substantial risk to the survival of humanity. While most research focuses on predicting annual CO2 emissions, which are crucial for setting long-term emission mitigation targets, the precise prediction of daily CO2 emissions is equally vital for setting short-term targets. This study examines the performance of 14 models in predicting daily CO2 emissions data from 1/1/2022 to 30/9/2023 across the top four polluting regions (China, India, the USA, and the EU27&UK). The 14 models used in the study include four statistical models (ARMA, ARIMA, SARMA, and SARIMA), three machine learning models (support vector machine (SVM), random forest (RF), and gradient boosting (GB)), and seven deep learning models (artificial neural network (ANN), recurrent neural network variations such as gated recurrent unit (GRU), long short-term memory (LSTM), bidirectional-LSTM (BILSTM), and three hybrid combinations of CNN-RNN). Performance evaluation employs four metrics (R2, MAE, RMSE, and MAPE). The results show that the machine learning (ML) and deep learning (DL) models, with higher R2 (0.714–0.932) and lower RMSE (0.480–0.247) values, respectively, outperformed the statistical model, which had R2 (− 0.060–0.719) and RMSE (1.695–0.537) values, in predicting daily CO2 emissions across all four regions. The performance of the ML and DL models was further enhanced by differencing, a technique that improves accuracy by ensuring stationarity and creating additional features and patterns from which the model can learn. Additionally, applying ensemble techniques such as bagging and voting improved the performance of the ML models by approximately 9.6%, whereas hybrid combinations of CNN-RNN enhanced the performance of the RNN models. In summary, the performance of both the ML and DL models was relatively similar. However, due to the high computational requirements associated with DL models, the recommended models for daily CO2 emission prediction are ML models using the ensemble technique of voting and bagging. This model can assist in accurately forecasting daily emissions, aiding authorities in setting targets for CO2 emission reduction.

  • Research Article
  • Cite Count Icon 5
  • 10.1016/j.ejrad.2025.112060
Predicting hepatocellular carcinoma response to TACE: A machine learning study based on 2.5D CT imaging and deep features analysis.
  • Jun 1, 2025
  • European journal of radiology
  • Chong Lin + 4 more

Predicting hepatocellular carcinoma response to TACE: A machine learning study based on 2.5D CT imaging and deep features analysis.

  • Book Chapter
  • Cite Count Icon 1
  • 10.1007/978-3-031-11349-9_6
Comparative Analysis of Machine Learning and Deep Learning Models for Ship Classification from Satellite Images
  • Jan 1, 2022
  • Abhinaba Hazarika + 2 more

The automatic detection of the ship from satellite image analysis is the limelight of research in recent years due to its widespread applications. In this paper, a handful of traditional machine learning and deep learning models are compared based on their performance to classify the satellite images available in the public repository as a ship or other categories. The Support Vector Machine(SVM), Decision Trees, Random Forest, K-Nearest Neighbor (KNN), Gaussian Naive Bayes (GaussianNB), and Logistic Regression are machine learning models used in the present work. Histogram of Gradient (HoG) features are used as feature descriptors considering the diverse size and shape of ships in the satellite image dataset. Transfer learning is applied using the deep learning models namely, Inception and ResNet, that are fine-tuned for various learning rates and optimizers. The meticulous experimentation carried out reveals that traditional machine learning performs well when trained and tested on a single dataset. However, there is a drastic change in the performance of machine learning models when tested on a different ship dataset. The results show that the deep learning models have better feature detection and thus have better performance when transfer learning is used on various datasets.KeywordsShip classificationMachine learningDeep learning

  • Book Chapter
  • 10.1039/9781837673179-00224
Predicting Solid-state NMR Observables via Machine Learning
  • Mar 31, 2025
  • Pablo A Unzueta + 1 more

Machine learning is becoming increasingly important in the prediction of nuclear magnetic resonance (NMR) chemical shifts and other observable properties. This chapter provides an introduction to the construction of machine learning (ML) models for predicting NMR properties, including the discussion of feature engineering, common ML model types, Δ-ML and transfer learning, and the curation of training and testing data. Then it discusses a number of recent examples of ML models for predicting chemical shifts and spin–spin coupling constants in organic and inorganic species. These examples highlight how the decisions made in constructing the ML model impact its performance, discuss strategies for achieving more accurate ML models, and present some representative case studies showing how ML is transforming the way NMR crystallography is performed.

  • Research Article
  • Cite Count Icon 108
  • 10.1039/c8nr05703f
Machine learning and artificial neural network prediction of interfacial thermal resistance between graphene and hexagonal boron nitride.
  • Jan 1, 2018
  • Nanoscale
  • Hong Yang + 3 more

High-performance thermal interface materials (TIMs) have attracted persistent attention for the design and development of miniaturized nanoelectronic devices; however, a large number of potential new materials exist to form these heterostructures and the explorations of their thermal properties are time consuming and expensive. In this work, we train several supervised machine learning (ML) and artificial neural network (ANN) models to predict the interfacial thermal resistance (R) between graphene and hexagonal boron-nitride (hBN) with only the knowledge of the system temperature, coupling strength between two layers, and in-plane tensile strains. The training data were obtained by high-throughput computations (HTCs) of R using classical molecular dynamics (MD) simulations. Four different ML models, i.e., linear regression, polynomial regression, decision tree and random forest, are explored. A pair of one dense layer ANNs and another pair of two dense layer deep neural networks (DNNs) are also investigated. It is reported that the DNN models provide better R prediction results compared to the ML models. The thermal property predictions using HTC and ML/ANN models are applicable to a wide range of materials and open up new perspectives in the explorations of TIMs.

  • Conference Article
  • Cite Count Icon 2
  • 10.2514/6.2023-4184
Multi-Fidelity Propeller Noise Prediction using a Data-Driven Approach
  • Jun 8, 2023
  • Beckett Yx Zhou + 5 more

Accurate and rapid prediction of propeller aeroacoustic performance is of paramount importance to the design of quiet urban air mobility (UAM) concepts emerging over the last decade. For the broadband noise component, Amiet-type low-fidelity methods are typically used in order to avoid resorting to computationally prohibitive scale-resolving simulations. These methods however, are highly limited in their accuracy and robustness over the range of flow conditions UAM propellers are designed to operate. This paper presents an exploratory effort in developing a multi-fidelity framework to enhance the propeller broadband noise prediction using data-driven approaches. In this framework, a deep neural network machine learning (ML) model is trained in a multi-fidelity manner using transfer learning (TL). The model is first trained using a large number of computationally inexpensive low-fidelity simulations and then enhanced by a small number of high-fidelity aeroacoustic wind tunnel measurements using TL. In the deployment stage, this ML model enables rapid prediction of broadband sound pressure level of isolated 2-bladed propellers between 1kHz and 7kHz at three farfield observer locations given the cross-sectional profile, pitch-to-diameter ratio as well as forward speed and rotational speed of the propeller. Results evaluated based on experimental data from two extra sets of propeller blades, which were withheld from the ML training process, indicated that the multi-fidelity ML model trained with both low- and high-fidelity data based on the TL approach is capable of significantly improving the predictive accuracy of the low-fidelity model. Furthermore, the TL-based multi-fidelity ML model delivered more accurate predictions than the ML models trained solely with limited high-fidelity data.

  • Research Article
  • Cite Count Icon 11
  • 10.1111/exsy.13153
Deep learning‐based smishing message identification using regular expression feature generation
  • Oct 5, 2022
  • Expert Systems
  • Aakanksha Sharaff + 2 more

The increase in the number of undesired SMS termed smishing message and the data imbalance problem has generated a great demand for the development of more reliable anti‐spam filters. State of the art machine learning approaches are being employed to recognize and separate spam messages. Most recent studies target message classification by using numerous properties and features of the words but fail to consider the circumstantial features like long‐range dependencies between the words that are extremely important in identifying smishing messages. The idea is to develop an intelligent model that will distinguish between smishing messages and ham messages, by adopting a combined approach of regular expression (Regex), machine learning (ML) and deep learning (DL) models. Regex rules are generated using the dataset's spam messages for the purpose of refining the dataset. Support vector machine (SVM), Multinomial Naive Bayes and Random Forest are included under machine learning models and long short‐term memory (LSTM), bidirectional long short‐term memory (Bi‐LSTM), stacked LSTM and stacked Bi‐LSTM are included under deep learning models. The comparison between machine learning models and deep learning models is also carried out based on the performance evaluation parameters namely accuracy, precision, recall and F1 score of the models. It is observed that deep learning models perform better than machine learning models and the introduction of a regular expression to the dataset increases the efficiency of both the deep learning models and machine learning models.

  • Conference Article
  • 10.1109/mlise57402.2022.00015
Transferability of Pretrained Convolutional Neural Networks for Breast Cancer Detection
  • Aug 1, 2022
  • Zeyu Jin + 2 more

Breast cancer is the most common cancer in the world. In breast cancer, the invasive ducal carcinoma (IDC) is the most common breast cancer. For the past few years, more and more machine learning models have been used in medical treatment as an auxiliary tool for doctors. In this work, the machine learning model mainly including Convolutional Neural Network (CNN) and transfer learning was employed to identify breast cancer images to test if the histopathological breast sections are IDC. Since the process of diagnosing IDC by naked eyes is tedious, and time-costy, it could be helpful to develop a model to let machine classify the pathological section, such that the process would be more efficient. While based on previous study, it has shown that deep learning methods could achieve high performance, but it remains to be a question, that if the pre-trained models trained on datasets with no weights on medical images could be helpful while applying on tasks such as this study. The goal is to explore how transfer learning performs compare to other methods. In previous breast cancer recognition work, the SVM model did a pretty good job, so, in this work, a comparison will be done among CNN, transfer learning and SVM. In this study, methods such as cross validation, feature scaling, Support Vector Machine (SVM), CNN and transfer learning were applied to find which one can Figure out the types of pathological sections in the best way; different approaches to evaluate the performance of SVM, CNN and transfer learning model were applied, and find the best one to classify breast cancer pathological sections; the best model were selected regarding to these standards. Finally, it turns out that pre-trained model DenseNet201 for transfer learning works very well. This suggests that transfer learning can achieve good performance in this case by utilizing the weight of ImageNet.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant