Predicting Bankruptcy at Polish Companies: A Comparison of Selected Machine Learning and Deep Learning Algorithms
This study compares machine learning and deep learning algorithms, including gradient boosting, random forest, and neural networks, for predicting bankruptcy among Polish companies from 2008–2017. Results show comparable efficiency across methods, with models trained on distinct datasets from different periods and companies.
Insolvency prediction is one of the crucial abilities in corporate finance and financial management. It is critical in accounts receivable management, capital budgeting decisions, financial analysis, capital structure management, going concern assessment and co-operation with other companies. The purpose of this paper is to compare the efficiency of selected deep learning and machine learning algorithms trained on a representative sample of Polish companies for the period 2008–2017. In particular, the paper tested the following popular machine learning algorithms: discriminant analysis (DA), logit (L), support vector machines (SVM), random forest (RF), gradient boosting decision trees (GB), neural network with one hidden layer (NN), convolutional neural network (CNN), and naïve Bayes (NB). The research hypotheses evaluated in the paper state that if one has access to a large sample of companies, the most accurate algorithm (first choice) in bankruptcy prediction will be gradient boosting decision trees (H1), random forest (H2) and neural networks (H3) (deep learning) algorithms. The initial hypotheses were formulated based on the practitioners’ opinions regarding the usefulness of various machine learning and artificial intelligence algorithms in bankruptcy prediction. As the results of the research suggest, both deep learning and machine learning algorithms proved to have very comparable efficiency. The new factor introduced in the paper was that the training of the models was carried out on a representative sample of companies (for years 2008–2013) and also the testing phase used a significant number of bankrupt and active companies (validation included a completely different set of companies than those used in the training phase: data were taken from a different time period, 2014–2017, and companies in both sets were also completely different).
- Book Chapter
1
- 10.1007/978-981-19-2821-5_59
- Sep 27, 2022
The main objective of this research is to analyze and compare the performance of machine learning (ML) and deep learning (DL) algorithms in detecting online hate speech. Therefore, Support Vector Machine (SVM), Random Forest (RF), Decision Tree (DT), Logistic Regression (LR), Convolution Neural Network (CNN), Recurrent Neural Network_Long Short-Term Memory (RNN_LSTM), BERT (Bidirectional Encoder Representations from Transformers), and Distil BERT algorithms have been explored and analyzed in this research. This research has applied the dataset on hate speech which was developed by Andry Samoshyn which is publicly available in Kaggle. ML algorithms and DL algorithms have got good scores in accuracy. In ML, SVM, RF, and LR have got top accuracy values. In DL algorithms, RNN_LSTM, Distil BERT, and BERT have performed well in accuracy. Based on F-measurement, DL classifiers have outperformed ML algorithms. Distil BERT has obtained the highest F-measurement scores. When we compare the overall performances, DL is performed well rather than ML in detecting hate speech. Especially transformer-based models of DL are more efficient than other DL and ML algorithms.KeywordsHate speechMachine learningDeep learning TwitterAnd performance comparison
- Research Article
15
- 10.1080/23279095.2024.2382823
- Jul 31, 2024
- Applied Neuropsychology: Adult
The cognitive impairment known as dementia affects millions of individuals throughout the globe. The use of machine learning (ML) and deep learning (DL) algorithms has shown great promise as a means of early identification and treatment of dementia. Dementias such as Alzheimer’s Dementia, frontotemporal dementia, Lewy body dementia, and vascular dementia are all discussed in this article, along with a literature review on using ML algorithms in their diagnosis. Different ML algorithms, such as support vector machines, artificial neural networks, decision trees, and random forests, are compared and contrasted, along with their benefits and drawbacks. As discussed in this article, accurate ML models may be achieved by carefully considering feature selection and data preparation. We also discuss how ML algorithms can predict disease progression and patient responses to therapy. However, overreliance on ML and DL technologies should be avoided without further proof. It’s important to note that these technologies are meant to assist in diagnosis but should not be used as the sole criteria for a final diagnosis. The research implies that ML algorithms may help increase the precision with which dementia is diagnosed, especially in its early stages. The efficacy of ML and DL algorithms in clinical contexts must be verified, and ethical issues around the use of personal data must be addressed, but this requires more study.
- Conference Article
- 10.1109/icecer65523.2025.11401299
- Dec 6, 2025
This study presents a comparative analysis of the classification performance of facial and emotion recognition systems using Machine Learning (ML) and Deep Learning (DL) algorithms. The primary objective of this work is to evaluate the applicability of emotion recognition in fields such as psychotherapy and crime analysis, using the FER-2013 dataset.The study was conducted with ML algorithms such as Support Vector Machines (SVM), Random Forest (RF), K-Nearest Neighbor (KNN), Decision Trees, and Gradient Boosting, as well as the DL algorithm, Convolutional Neural Network (CNN). Supported by the application of feature selection, preprocessing, and feature extraction techniques, the model’s performance was measured using standard metrics such as accuracy, precision, F1-score, and AUC-ROC.The experimental results showed that the highest classification accuracy (46.61%) was achieved with the CNN model. While the ML models generally offered lower accuracy, they provided advantages in terms of computational efficiency in specific scenarios.This study aims to contribute to the literature by providing a comparative analysis of ML and DL algorithms and by highlighting the effect of data preprocessing on performance. The findings set targets for future work, such as real-time system integration and hybrid model development.
- Research Article
- 10.26562/ijirae.2025.v1212.02
- Dec 11, 2025
- International Journal of Innovative Research in Advanced Engineering
The proliferation of Android malware has become a significant concern in the cyber security landscape. Traditional signature-based detection methods are no longer effective against the rapidly evolving malware threats. Machine Learning (ML) and Deep Learning (DL) algorithms have emerged as a promising solution for Android malware classification and detection. This study aims to investigate the effectiveness of various ML and DL algorithms for Android malware detection. This work analyses the performance of several algorithms, including Support Vector Machines (SVM), Random Forest (RF), and Recurrent Neural Network (RNN). From this analysis, reveals the DL algorithms, particularly RNN, outperform traditional ML algorithms in terms of accuracy, precision, and recall. This finding suggests that the DL algorithm and selects the features can provide an effective solution for Android malware classification and detection. The results of this study can be used to develop a robust and efficient Android malware detection system, which can help protect against the increasing threats of mobile malware. Overall, this study demonstrates the potential of ML and DL algorithms in detecting Android malware and provides insights into the development of effective detection systems.
- Research Article
- 10.55041/ijsrem27894
- Jan 4, 2024
- INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
Wind energy has the possibility for bringing out energy in a very constant and sustainable manner being a notable and eligible source. Although, wind energy does include numerous challenges like, the halted asset of wind plants, early investment costs, and therefore, the strain in discovering areas of wind efficiency. The major objective for proposing this work is to determine the power efficiency of wind turbines, which will also aid in the formulation of a proposal to reduce wind turbine maintenance costs. During this research, data analysis of turbine generators is performed on day to day wind speed info using machine learning and deep learning algorithms. A way is put forward by us to support deep learning and machine learning algorithms which can predict different values of power reliably. Hence, the execution of machine and deep learning algorithms are analyzed. For forecasting for a longer term these algorithms may be used for wind generation rate with historical relation to wind speed info. Index Terms: Wind turbine,machine learning algorithm
- Research Article
118
- 10.1007/s10661-024-12454-z
- Feb 24, 2024
- Environmental Monitoring and Assessment
Digital image processing has witnessed a significant transformation, owing to the adoption of deep learning (DL) algorithms, which have proven to be vastly superior to conventional methods for crop detection. These DL algorithms have recently found successful applications across various domains, translating input data, such as images of afflicted plants, into valuable insights, like the identification of specific crop diseases. This innovation has spurred the development of cutting-edge techniques for early detection and diagnosis of crop diseases, leveraging tools such as convolutional neural networks (CNN), K-nearest neighbour (KNN), support vector machines (SVM), and artificial neural networks (ANN). This paper offers an all-encompassing exploration of the contemporary literature on methods for diagnosing, categorizing, and gauging the severity of crop diseases. The review examines the performance analysis of the latest machine learning (ML) and DL techniques outlined in these studies. It also scrutinizes the methodologies and datasets and outlines the prevalent recommendations and identified gaps within different research investigations. As a conclusion, the review offers insights into potential solutions and outlines the direction for future research in this field. The review underscores that while most studies have concentrated on traditional ML algorithms and CNN, there has been a noticeable dearth of focus on emerging DL algorithms like capsule neural networks and vision transformers. Furthermore, it sheds light on the fact that several datasets employed for training and evaluating DL models have been tailored to suit specific crop types, emphasizing the pressing need for a comprehensive and expansive image dataset encompassing a wider array of crop varieties. Moreover, the survey draws attention to the prevailing trend where the majority of research endeavours have concentrated on individual plant diseases, ML, or DL algorithms. In light of this, it advocates for the development of a unified framework that harnesses an ensemble of ML and DL algorithms to address the complexities of multiple plant diseases effectively.
- Research Article
32
- 10.1016/j.compeleceng.2023.108691
- Mar 22, 2023
- Computers and Electrical Engineering
Nine novel ensemble models for solar radiation forecasting in Indian cities based on VMD and DWT integration with the machine and deep learning algorithms
- Research Article
13
- 10.31083/j.rcm2501008
- Jan 8, 2024
- Reviews in cardiovascular medicine
Atrial fibrillation (AF) is a common arrhythmia that can result in adverse cardiovascular outcomes but is often difficult to detect. The use of machine learning (ML) algorithms for detecting AF has become increasingly prevalent in recent years. This study aims to systematically evaluate and summarize the overall diagnostic accuracy of the ML algorithms in detecting AF in electrocardiogram (ECG) signals. The searched databases included PubMed, Web of Science, Embase, and Google Scholar. The selected studies were subjected to a meta-analysis of diagnostic accuracy to synthesize the sensitivity and specificity. A total of 14 studies were included, and the forest plot of the meta-analysis showed that the pooled sensitivity and specificity were 97% (95% confidence interval [CI]: 0.94-0.99) and 97% (95% CI: 0.95-0.99), respectively. Compared to traditional machine learning (TML) algorithms (sensitivity: 91.5%), deep learning (DL) algorithms (sensitivity: 98.1%) showed superior performance. Using multiple datasets and public datasets alone or in combination demonstrated slightly better performance than using a single dataset and proprietary datasets. ML algorithms are effective for detecting AF from ECGs. DL algorithms, particularly those based on convolutional neural networks (CNN), demonstrate superior performance in AF detection compared to TML algorithms. The integration of ML algorithms can help wearable devices diagnose AF earlier.
- Conference Article
1
- 10.1109/icscds53736.2022.9760818
- Apr 7, 2022
Wind energy being a notable and eligible source, has the possibility for bringing out energy in a very constant and sustainable manner. However, wind energy does include numerous challenges like, the halted asset of wind plants, early investment costs, and the strain in discovering areas of wind efficiency. The major objective for proposing this work is to determine the power efficiency of wind turbines, which also aids in the formulation of a proposal to reduce wind turbine maintenance costs. During this research, data analysis of turbine generators is performed on day-to-day wind speed info using machine learning and deep learning algorithms. A way is put forward to support deep learning and machine learning algorithms which can predict different values of power reliably. Hence, the execution of machine and deep learning algorithms are analyzed. For forecasting for a longer term, these algorithms may be used for wind generation rate with historical relation to wind speed info. Moreover, the application of deep and machine learning-based models is place distinct to that of model-trained places. This data analysis demonstrates that in unspecified geographies of wind plants, these sets of algorithms could be successfully implied by utilizing the base location model. The entire project focuses on wind turbine generators and includes the use of data visualization of data analytics to analyze the data and detect the factors that influence wind power generation. With the support of previous data output, wind power is anticipated using both machine learning and deep learning models, where different datasets are used for training and testing. This adds to the uniqueness of this work.
- Research Article
1
- 10.52783/jisem.v10i7s.857
- Jan 27, 2025
- Journal of Information Systems Engineering and Management
This study offers an in-depth comparative examination of predictive monitoring systems utilizing sophisticated machine learning (ML) and deep learning (DL) algorithms. The investigation delves into the efficacy, advantages, and constraints of ML algorithms, exemplified by Random Forest and Extreme Gradient Boosting will be contrasted with Deep Learning (DL) algorithms including artificial neural networks (ANNs) and LSTM, recurrent neural networks (RNNs) in this study. By evaluating these algorithms across diverse domains, the research aims to discern optimal strategies for predictive monitoring, considering factors like efficiency, real-time processing, and adaptability. The findings contribute valuable insights for practitioners and researchers, informing the selection and deployment of algorithms in predictive monitoring systems
- Dissertation
2
- 10.31274/td-20240617-187
- Jan 1, 2024
In the evolving landscape of cyber-physical systems (CPS), such as Electric Vehicle Charging Stations (EVCS) and the Industrial Internet of Things (IIoT), the convergence of cyber and physical domains introduces a myriad of opportunities for enhanced efficiency and connectivity. However, this integration also presents substantial security challenges, with vulnerabilities posing risks to both cyber operations and physical system functionality. Anomalies, indicative of cyber-attacks or system malfunctions, manifest as deviations from established operational norms, necessitating sophisticated detection mechanisms. This thesis presents a comprehensive investigation into the application of machine learning (ML) and deep learning (DL) algorithms for anomaly detection within these critical CPS frameworks, emphasizing IIoT and EVCS systems. Acknowledging the inherent complexity of these systems, the research initially applies a suite of ML algorithms—Support Vector Machines (SVM), Decision Trees (DT), and Random Forests (RF)—to IIoT systems, exploiting the relatively straightforward operational patterns to establish a foundational anomaly detection framework. This strategic application leverages the diverse strengths of each algorithm: SVM’s capacity for handling high-dimensional data, DT’s interpretability and ease of use, and RF’s robustness and accuracy in classification tasks. Subsequently, the thesis escalates the analytical depth by incorporating Long Short-Term Memory (LSTM) networks, a DL-based technique, to navigate the more intricate anomaly detection challenges encountered in both IIoT and EVCS systems. LSTM networks are selected for their proven efficacy in processing and making predictions based on long-term dependencies in time-series data, a common characteristic in the operational data of EVCS and IIoT systems. This transition underscores a methodological advancement towards models capable of capturing complex, temporal data relationships, essential for detecting sophisticated anomalies. We carried out ML-based anomaly detection using the WUSTL-IIoT datasets from 2018 and 2021. The ML algorithms underwent extensive training and evaluation, demonstrating substantial effectiveness. Specifically, the SVM and DT models attained an accuracy of 97.6%, with the RF model achieving a slightly superior accuracy of 98.8%. To enhance the detection capability, an LSTM model was implemented, which achieved a remarkable accuracy rate of 99.57%. This performance exemplifies the advanced potential of DL methodologies in navigating the complexities of anomaly detection within intricate data environments. We carried out LSTM-based anomaly detection in EVCS systems by utilizing the CICEVCS 2023 and 2024 datasets. These datasets, encompassing a wide range of attack scenarios along with normal operational data, provided a complex backdrop for the application of the LSTM model. The DL algorithm skillfully navigated these complexities, achieving an impressive accuracy rate of 99.589% in identifying anomalies. This experiment underscores the advanced capabilities of DL, specifically LSTM, in accurately analyzing and predicting anomalies across comprehensive time-series data streams within EVCS systems. The findings from these case studies highlight the pivotal role of ML and DL algorithms in advancing anomaly detection capabilities within IIoT and EVCS systems. By meticulously applying and evaluating SVM, DT, RF, and LSTM models against real-world operational and attack scenarios, this thesis demonstrates the efficacy of these computational techniques in identifying anomalies and enhances the strategic framework for securing CPS against emerging cyber threats.
- Research Article
67
- 10.3390/jpm10040286
- Dec 16, 2020
- Journal of Personalized Medicine
Brain magnetic resonance imaging (MRI) is useful for predicting the outcome of patients with acute ischemic stroke (AIS). Although deep learning (DL) using brain MRI with certain image biomarkers has shown satisfactory results in predicting poor outcomes, no study has assessed the usefulness of natural language processing (NLP)-based machine learning (ML) algorithms using brain MRI free-text reports of AIS patients. Therefore, we aimed to assess whether NLP-based ML algorithms using brain MRI text reports could predict poor outcomes in AIS patients. This study included only English text reports of brain MRIs examined during admission of AIS patients. Poor outcome was defined as a modified Rankin Scale score of 3–6, and the data were captured by trained nurses and physicians. We only included MRI text report of the first MRI scan during the admission. The text dataset was randomly divided into a training and test dataset with a 7:3 ratio. Text was vectorized to word, sentence, and document levels. In the word level approach, which did not consider the sequence of words, and the “bag-of-words” model was used to reflect the number of repetitions of text token. The “sent2vec” method was used in the sensation-level approach considering the sequence of words, and the word embedding was used in the document level approach. In addition to conventional ML algorithms, DL algorithms such as the convolutional neural network (CNN), long short-term memory, and multilayer perceptron were used to predict poor outcomes using 5-fold cross-validation and grid search techniques. The performance of each ML classifier was compared with the area under the receiver operating characteristic (AUROC) curve. Among 1840 subjects with AIS, 645 patients (35.1%) had a poor outcome 3 months after the stroke onset. Random forest was the best classifier (0.782 of AUROC) using a word-level approach. Overall, the document-level approach exhibited better performance than did the word- or sentence-level approaches. Among all the ML classifiers, the multi-CNN algorithm demonstrated the best classification performance (0.805), followed by the CNN (0.799) algorithm. When predicting future clinical outcomes using NLP-based ML of radiology free-text reports of brain MRI, DL algorithms showed superior performance over the other ML algorithms. In particular, the prediction of poor outcomes in document-level NLP DL was improved more by multi-CNN and CNN than by recurrent neural network-based algorithms. NLP-based DL algorithms can be used as an important digital marker for unstructured electronic health record data DL prediction.
- Research Article
8
- 10.5194/isprs-archives-xlviii-g-2025-263-2025
- Jul 28, 2025
- The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Abstract. An updated map of the area's land use land cover (LULC) is necessary for strategic planning and management of land use to shape the town sustainably. The advances in remote sensing imageries and artificial intelligence have facilitated the extraction of LULC classification. With the high number of studies on LULC mapping using various machine learning (ML) and deep learning (DL) algorithms incorporating imageries, no established algorithm shows stable results for all the datasets and study regions. Therefore, we used three robust machine learning algorithms, Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbour (KNN), and four deep learning algorithms, Residual Network (ResNet50 and ResNet152) and Visual Geometry Group (VGG16 and VGG19), to understand which model can produce a highly accurate LULC map in the Indian context, which are inherently unplanned and unorganized using Sentinel 2 imageries. The results of these models were then comparatively analyzed statistically using Accuracy, Recall, Precision, F1-score, and Kappa coefficient. Although DL models require a large number of training datasets, they outperformed the ML algorithms with higher Kappa coefficient values (ResNET50 = 0.90, ResNET-152 = 0.91, VGG-16 = 0.94, VGG-19 = 0.94). VGG-19 has consistently given better performance in all accuracy metrics. Overall the study highlights the potential of deep learning models, particularly VGG-19, in generating highly accurate LULC maps for complex and unplanned urban environments in India. These findings underscore the importance of leveraging advanced AI techniques in remote sensing for effective land use planning and sustainable urban development.
- Preprint Article
- 10.5194/egusphere-egu23-15915
- May 15, 2023
Ground deformation caused by groundwater exploitation leads to significant socio-economic losses worldwide. Driving factors such as population growth and climate change will increase these losses, especially in arid regions where droughts are becoming more intense, longer lasting, and frequent. Therefore, there is a need to generate models capable of forecasting ground deformation. However, few studies have analyzed deformation time series (DTS) to identify and characterize subsidence phenomena.Our research aims to predict the ground deformation associated with groundwater abstractions in 18 wells of the Madrid Detrital Aquifer (ATDM) using statistical models and shallow and deep Machine Learning (ML) algorithms. We generated a database with 18 monthly time series (one for each well) between 1992 and 2010, with data for two variables: a binary variable indicating extraction-recovery cycles of the aquifer and a continuous variable representing the average deformation for the area of influence of each well. DTS generated from Persistent Scatter Interferometry (PSI) of ERS-1/2 and ENVISAT radar images were used to calculate the average deformation. Finally, we applied six different methods for forecasting DTS: two statistical models, Autoregressive Integrated Moving Average (ARIMA) and Prophet (P), one ensemble shallow ML algorithm, Random Forest (RF), one hybrid method, Neural Prophet (NP), and two Deep Learning (DL) techniques 1D Convolutional Neural Networks (CNN1D), and Long Short-Term Memory (LSTM).The analysis of DTS allowed us to differentiate two zones with different hydrological behavior: a zone of higher permeability (north zone) and another of lower permeability (south zone). We found that establishing the architectures of ML and DL algorithms based on hydrological zones improves the prediction of ground deformation. ML and DL algorithms provide better forecasts compared to statistical and hybrid models. Specifically, LSTM and RF offer the best results. Our results show the potential of LSTM algorithms and the previous grouping of DTS in predicting ground deformation associated with groundwater exploitation.This work has been developed thanks to the pre-doctoral grant for the Training of Research Personnel (PRE2021-100044) funded by MCIN/AEI/10.13039/501100011033 and by "FSE invests in your future" within the framework of the SARAI project "Towards a smart exploitation of land displacement data for the prevention and mitigation of geological-geotechnical risks" PID2020-116540RB-C22 funded by MCIN/AEI/10.13039/501100011033.
- Research Article
53
- 10.1016/j.cose.2023.103143
- Feb 17, 2023
- Computers & Security
Feature mining for encrypted malicious traffic detection with deep learning and other machine learning algorithms