Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Garbage In, Garbage Out? The Impact of Data Quality on the Performance of Financial Distress Prediction Models

  • TL;DR
  • Abstract
  • Literature Map
  • Similar Papers
TL;DR

This study demonstrates that structured data preparation significantly enhances financial distress prediction accuracy, with an average improvement of 15.6 percentage points across models, especially for decision trees and neural networks, highlighting data quality as a key factor alongside algorithm choice.

Abstract
Translate article icon Translate Article Star icon

Financial distress prediction remains a central topic in corporate finance and risk management, with extensive research devoted to improving classification accuracy through increasingly sophisticated statistical and machine learning techniques. Nevertheless, the influence of data preparation on predictive performance has received comparatively less systematic attention. This study examines how an economically grounded data-preparation process affects the predictive performance of selected statistical and machine-learning models dedicated to predicting corporate financial distress. Using the chosen financial ratios, generally accepted indicators of corporate financial stability and economic performance, financial distress models are estimated on both raw, unprocessed input data and pre-processed data involving the exclusion of economically implausible accounting values, treatment of missing observations, and class balancing. In light of the above, the study adopts a structured methodological approach to assess the predictive performance of selected classification models, namely decision tree algorithms (CART, CHAID, and C5.0), artificial neural networks (ANNs), logistic regression (LR), and linear discriminant analysis (DA), using confusion-matrix–based evaluation and a comprehensive set of evaluation measures. The results suggest that the process of input data preparation is a critical factor, significantly improving the predictive performance of financial distress prediction models across most modelling techniques employed. The most pronounced gains are observed in decision tree models. ANNs also demonstrate marked improvement after input data preparation, whereas LR benefits more moderately, and linear DA remains limited despite preprocessing. The average gain in accuracy across all six modelling techniques, calculated as the difference between pre-processed and raw performance for each method and averaged across methods, was approximately 15.6 percentage points, with specificity improving by approximately 26.9 percentage points on average, amounting to roughly half the performance variation attributable to algorithm choice, which underscores that data preparation is a primary determinant of model reliability alongside algorithm selection. A step-level detailed analysis further shows that missing value imputation is the dominant driver of improvement for tree-based models, while class balancing contributes most for ANNs and logistic regression. The findings highlight that reliable financial distress prediction depends not only on technique selection but also on the consistency and economic plausibility of the input data, underscoring the central role of structured data preparation in developing robust early-warning models.

Similar Papers
  • Research Article
  • Cite Count Icon 274
  • 10.1016/j.eswa.2008.03.020
Using neural networks and data mining techniques for the financial distress prediction model
  • Apr 15, 2008
  • Expert Systems with Applications
  • Wei-Sen Chen + 1 more

Using neural networks and data mining techniques for the financial distress prediction model

  • Research Article
  • Cite Count Icon 11
  • 10.3934/era.2023240
A deep learning approach of financial distress recognition combining text
  • Jan 1, 2023
  • Electronic Research Archive
  • Jiawang Li + 1 more

<abstract><p>The financial distress of listed companies not only harms the interests of internal managers and employees but also brings considerable risks to external investors and other stakeholders. Therefore, it is crucial to construct an efficient financial distress prediction model. However, most existing studies use financial indicators or text features without contextual information to predict financial distress and fail to extract critical details disclosed in Chinese long texts for research. This research introduces an attention mechanism into the deep learning text classification model to deal with the classification of Chinese long text sequences. We combine the financial data and management discussion and analysis Chinese text data in the annual reports of 1642 listed companies in China from 2017 to 2020 in the model and compare the effects of the data on different models. The empirical results show that the performance of deep learning models in financial distress prediction overcomes traditional machine learning models. The addition of the attention mechanism improved the effectiveness of the deep learning model in financial distress prediction. Among the models constructed in this study, the Bi-LSTM+Attention model achieves the best performance in financial distress prediction.</p></abstract>

  • Research Article
  • Cite Count Icon 172
  • 10.1561/1400000018
Financial Statement Analysis and the Prediction of Financial Distress
  • May 17, 2011
  • Foundations and Trends® in Accounting
  • William H Beaver + 2 more

Financial statement analysis has been used to assess a company’s likelihood of financial distress – the probability that it will not be able to repay its debts. Financial statement analysis was used by credit suppliers to assess the credit worthiness of its borrowers. Today, financial statement analysis is ubiquitous and involves a wide variety of ratios and a wide variety of users, including trade suppliers, banks, creditrating agencies, investors and management, among others. Financial distress refers to the inability of a company to pay its financial obligations as they mature. Empirically, academic research in accounting and finance has focused on either bond default or bankruptcy. The basic issue is whether the probability of distress varies in a significant manner conditional upon the magnitude of the financial statement ratios. This monograph discusses the evolution of three main streams within the financial distress prediction literature: The set of dependent and explanatory variables used, the statistical methods of estimation, and the modeling of financial distress. For over 100 years, financial statement analysis has been used to assess a company’s likelihood of financial distress – the probability that it will not be able to repay its debts. Financial statement analysis was used by credit suppliers to assess the credit worthiness of its borrowers. In many cases, there was little alternative, reliable information, other than the general reputation of the borrower. A major force for the audit of financial statements arose from the demand to help ensure more reliable financial statements. For example, major users were trade suppliers allowing companies to purchase inventory on credit until the goods could be resold. For these users, there was an emphasis on shortterm ability to repay, given the focus on ability to repay over the period of inventory turnover (typically a matter of 30-60 days). In this context, the current ratio (the ratio of current assets to current liabilities) was one of the first and most prominent ratios used.1 Today, financial statement analysis is ubiquitous and involves a wide variety of ratios and a wide variety of users, including trade suppliers, banks, credit-rating agencies, investors and management, among others. Moreover, financial statements are only one among many sources of information about a company. Financial distress refers to the inability of a company to pay its financial obligations as they mature. Empirically, academic research in accounting and finance has focused on either bond default or bankruptcy. The basic issue is whether the probability of distress varies in a significant manner conditional upon the magnitude of the financial statement ratios. This monograph discusses the evolution of three main streams within the financial distress prediction literature: the set of dependent and explanatory variables used, the statistical methods of estimation, and the modeling of financial distress. The outline of the monograph is as follows: Section 1 discusses concepts of financial distress. Section 2 discusses theories regarding the use of financial ratios as predictors of financial distress. Section 3contains a brief review of the literature. Section 4 discusses the use of market price-based models of financial distress. Section 5 develops the statistical methods for empirical estimation of the probability of financial distress. Section 6 discusses the major empirical findings with respect to prediction of financial distress. Section 7 briefly summarizes some of the more relevant literature with respect to bond ratings.Section 8 presents some suggestions for future research, and Section 9 presents concluding remarks.

  • Conference Article
  • Cite Count Icon 1
  • 10.2991/jcis.2006.150
An Application of Intellectual Capital on Financial Distress Models by Using Neural Network
  • Jan 1, 2006
  • Kuang-Hua Hsu + 2 more

As the era of knowledge economy is prevalent in U.S. during 1992, knowledge economy plays an important role around the world. The value and competition of the traditional companies accounted on tangible assets. However, in the era of knowledge economy, the value and continuing operation of the companies accounted on intangible assets. It is not sufficient in estimating the value of a company only by financial ratios (Bublita and Ettredge, 1989; Chauvin and Hirschey, 1993; Bontis et al., 2000). The previous researches found there is a close relationship between the intangible assets and the value of a firm. In this paper, a set of financial ratios, corporate governance variables and intellectual capital indicators will be investigated in a financial distress prediction by employing the Logit model and neural network model. The conclusion in this paper is that the prediction of the financial distress models is more accurate for one year prior to failure. The financial distress can be accurately predicted up to 89.2% or 91.53 with the accuracy diminishing after the first year. The finding is the same with that in previous literature. The profitability of a firm is the most important factor in the earlier stage. The firm with a poor profitability may have a financial distress in the short run. However, in the long run, the operation administration (such as account receivable) and the intellectual capital (such as patent and R&D expenditure) have a significant impact on the financial situation of a firm.

  • Conference Article
  • Cite Count Icon 7
  • 10.2991/assehr.k.200529.084
Accuracy of Financial Distress Model Prediction: The Implementation of Artificial Neural Network, Logistic Regression, and Discriminant Analysis
  • Jan 1, 2020
  • Proceedings of the 1st Borobudur International Symposium on Humanities, Economics and Social Sciences (BIS-HESS 2019)
  • Triasesiarta Nur + 1 more

The ability to predict financial failure forms an essential topic in financial research. The various models developed to predict the occurrence of Financial Distress and serve as an early warning system for the company's stakeholders before bankruptcy occurs. Enhanced accuracy of the predictions improves the ability to mitigate its adverse effect. This study aims to build Financial Distress models using Artificial Neural Network Model, Logistic Regression, and Discriminant Analysis, based on samples taken from manufacture sectors in the Indonesia Stock Exchange in the period 2015-2018. Accuracy of the three techniques in predicting Financial Distress are compared and results indicate Artificial Neural Network Model gave a better performance than the other techniques. It is crucial to consider the choice of predictor variables that determined the success of the financial distress model.

  • Research Article
  • 10.47191/jefms/v9-i2-33
Financial Distress Prediction in Kenyan Saccos: A Comparative Analysis of Machine Learning Models
  • Feb 25, 2026
  • Journal of Economics, Finance And Management Studies
  • Dr Dennis Mucee Ncurai, Phd + 1 more

Financial distress prediction is crucial for assessing the economic health of any organization or nation; as a whole. This research focuses on developing a robust machine learning model to predict financial distress in Kenyan Savings and Credit Cooperative Organizations (SACCOs). Notably, this is one of the first studies to explore financial distress prediction specifically within the context of Kenyan SACCOs. The specific objectives included the assessment of the performance of machine learning predictive models using both financial and non-financial variables, the identification of predictors that significantly affect financial distress of Kenyan SACCOs, and determining which features are the most predictive. The study evaluates the effectiveness of various machine learning algorithms, including Decision Tree, K-Nearest Neighbors, Logistic Regression, Artificial Neural Networks, Naive Bayes Classifier, Random Forest, and Support Vector Machine, through nine performance evaluation metrics, with the ROC-AUC score as the primary measure. Results indicate that while K-Nearest Neighbors achieved an ROC-AUC score of 0.78, an ensemble model using Stochastic Gradient Boosting reached 0.81.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 55
  • 10.1108/aea-10-2019-0039
Prediction of financial distress in the Spanish banking system
  • Nov 21, 2019
  • Applied Economic Analysis
  • Jessica Paule-Vianez + 2 more

Purpose The purpose of this study is to construct the first short-term financial distress prediction model for the Spanish banking sector. Design/methodology/approach The concept of financial distress covers a range of different types of financial problems, in addition to bankruptcy, which is not common in the sector. The methodology used to predict financial problems was artificial neural networks using traditional financial variables according to the capital, assets, management, earnings, liquidity and sensibility system, as well as a series of macroeconomic variables, the impact of which has been proven in a number of studies. Findings The results obtained show that artificial neural networks are a highly suitable method for studying financial distress in Spanish credit institutions and for predicting all cases in which an entity has short-term financial problems. Originality/value This is the first work that tries to build a model of artificial neural networks to predict the financial distress in the Spanish banking system, grouping under the concept of financial distress, apart from bankruptcy, other financial problems that affect the viability of these entities.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 14
  • 10.24136/eq.2023.035
Artificial neural network and decision tree-based modelling of non-prosperity of companies
  • Dec 30, 2023
  • Equilibrium. Quarterly Journal of Economics and Economic Policy
  • Marek Durica + 2 more

Research background: Financial distress or non-prosperity prediction has been a widely discussed topic for several decades. Early detection of impending financial problems of the company is crucial for effective risk management and important for all entities involved in the company’s business activities. In this way, it is possible to take the actions in the management of the company and eliminate possible undesirable consequences of these problems. Purpose of the article: This article aims to innovate financial distress prediction through the creation of individual models and ensembles, combining machine learning techniques such as decision trees and neural networks. These models are developed using real data. Beyond serving as an autonomous and universal tool especially useful in the Slovak economic conditions, these models can also represent a benchmark for Central European economies confronting similar economic dynamics. Methods: The prediction models are created using a dataset consisting of more than 20 financial ratios of more than 19 thousand real companies. Partial models are created employing machine learning algorithms, namely decision trees and neural networks. Finally, all models are compared based on a wide range of selected performance metrics. During this process, we strictly use a data mining methodology CRISP-DM. Findings & value added: The research contributes to the evolution of financial prediction and reveals the effectiveness of ensemble modelling in predicting financial distress, achieving an overall predictive ability of nearly 90 percent. Beyond its Slovak origins, this study provides a framework for early financial distress prediction. Although the models are created for diverse industries within the Slovak economy, they could also be useful beyond national borders. Moreover, the CRISP-DM methodological framework enables its adaptability for companies in other countries.

  • Research Article
  • Cite Count Icon 3
  • 10.9734/ajeba/2022/v22i24906
Financial Distress Prediction: A Hybrid Tracking Model Approach
  • Dec 21, 2022
  • Asian Journal of Economics, Business and Accounting
  • Zong-De Shen + 1 more

The purpose of this study was to build a highly accurate corporate financial distress tracking and prediction model based on hybrid machine learning technology. The research data were from Taiwan Economic Journal, and the research subjects were enterprises with financial distress risk announced in September 2022. In consideration of enterprise features, this study excluded the finance and insurance industries. The research period was three years (2019, 2020, and 2021) before the distress announcement. This study matched enterprises with financial distress and enterprises without financial distress (normal enterprises) at a ratio of 1:1 for each year. The sample size for each year included 374 enterprises with financial distress and 374 enterprises without financial distress. This study applied several machine learning technologies. At first, important variables were screened by applying artificial neural networks (ANNs). Next, prediction models were built based on decision tree C5.0 and random forest (RF) and were compared. According to the empirical result, the ANN-RF model provided a higher accuracy.

  • Research Article
  • Cite Count Icon 88
  • 10.17578/3-2-1
A Multicriteria Discrimination Method for the Prediction of Financial Distress: The Case of Greece
  • Jun 1, 1999
  • Multinational Finance Journal
  • Michael Doumpos + 1 more

Financial distress prediction is an essential issue in finance. Especially in emerging economies, predicting the future financial situation of individual corporate entities is even more significant, bearing in mind the general economic turmoil that can be caused by business failures. The research on developing quantitative financial distress prediction models has been focused on building discriminant models distinguishing healthy firms from financially distressed ones. Following this discrimination approach, this paper explores the applicability of a new non–parametric multicriteria decision aid discrimination method, called M.H.DIS, to predict financial distress using data concerning the case of Greece. A comparison with discriminant and logit analysis is performed using both a basic and a holdout sample. The results show that M.H.DIS can be considered as a new alternative tool for financial distress prediction. Its performance is superior to discriminant analysis and comparable to logit analysis.

  • Research Article
  • 10.35912/jomaps.v2i3.2389
Prediction of financial distress in transportation and logistics companies before, during and after the Covid-19 pandemic listed on the Indonesia Stock Exchange
  • Sep 2, 2024
  • Journal of Multidisciplinary Academic and Practice Studies
  • Sinta Dewi

Purpose: This study aims to predict financial distress in transportation and logistics companies before, during, and after the Covid-19 pandemic. Research Methodology: The research subjects were 23 companies, and the study utilized an artificial neural network model. This study uses financial ratios, including the debt-to-asset ratio (DAR), Current Ratio (CR), and Return on Assets (ROA) as input variables in the artificial neural network architecture. The objectives of this study are to calculate the three ratios used as test data, determine the differences in financial ratios between companies reported as financially distressed and those not experiencing financial distress in the training data, identify the artificial neural network model architecture that produces good performance on the training data sample for use in testing predictions, and predict financial distress using an artificial neural network on transportation and logistics companies listed on the Indonesia Stock Exchange, which are part of the research sample. The research sample consisted of 20 transportation and logistics companies listed on the Indonesia Stock Exchange from 2015 to 2019. Results: The results reveal that companies reported as financially distressed have lower average values for the three ratios compared to companies not experiencing financial distress, making them suitable input variables. The best artificial neural network architecture in this study included an input layer with 60 neurons, a hidden layer with 15 neurons, and an output layer with one neuron. This architecture achieved a training performance mean square error (MSE) of 0.125004 and an R value of 50.00%. The study's findings suggest that 12 companies are predicted to experience financial distress.

  • Research Article
  • 10.35912/jomaps.v2i3.2352
Prediction of Financial Distress in Transportation and Logistics Companies before, During and after the Covid-19 Pandemic Listed on the Indonesia Stock Exchange
  • Aug 8, 2024
  • Journal of Multidisciplinary Academic and Practice Studies
  • Sinta Dewi + 1 more

Purpose: This study aims to predict financial distress in transportation and logistics companies before, during, and after the Covid-19 pandemic. Research Methodology: The research subjects were 23 companies, and the study utilized an artificial neural network model. This study uses financial ratios, including the debt-to-asset ratio (DAR), Current Ratio (CR), and Return on Assets (ROA) as input variables in the artificial neural network architecture. The objectives of this study are to calculate the three ratios used as test data, determine the differences in financial ratios between companies reported as financially distressed and those not experiencing financial distress in the training data, identify the artificial neural network model architecture that produces good performance on the training data sample for use in testing predictions, and predict financial distress using an artificial neural network on transportation and logistics companies listed on the Indonesia Stock Exchange, which are part of the research sample. The research sample consisted of 20 transportation and logistics companies listed on the Indonesia Stock Exchange from 2015 to 2019. Results: The results reveal that companies reported as financially distressed have lower average values for the three ratios compared to companies not experiencing financial distress, making them suitable input variables. The best artificial neural network architecture in this study included an input layer with 60 neurons, a hidden layer with 15 neurons, and an output layer with one neuron. This architecture achieved a training performance mean square error (MSE) of 0.125004 and an R value of 50.00%. The study's findings suggest that 12 companies are predicted to experience financial distress.

  • Research Article
  • Cite Count Icon 39
  • 10.1016/j.mlwa.2024.100527
Survey, classification and critical analysis of the literature on corporate bankruptcy and financial distress prediction
  • Jan 11, 2024
  • Machine Learning with Applications
  • Jinxian Zhao + 2 more

Survey, classification and critical analysis of the literature on corporate bankruptcy and financial distress prediction

  • Research Article
  • Cite Count Icon 200
  • 10.1016/j.irfa.2017.02.004
Financial distress prediction: The case of French small and medium-sized firms
  • Feb 12, 2017
  • International Review of Financial Analysis
  • Nada Mselmi + 2 more

Financial distress prediction: The case of French small and medium-sized firms

  • Research Article
  • Cite Count Icon 75
  • 10.1016/j.eswa.2012.02.058
Predicting financial distress of the South Korean manufacturing industries
  • Feb 16, 2012
  • Expert Systems with Applications
  • Jae Kwon Bae

Predicting financial distress of the South Korean manufacturing industries

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant