Goodness-of-fit tests or measuring contamination fractions?
Goodness-of-fit tests or measuring contamination fractions?
- Research Article
1
- 10.2308/isys-10247
- Mar 1, 2012
- Journal of Information Systems
Book Review
- Research Article
56
- 10.1016/j.accinf.2018.09.004
- Nov 13, 2018
- International Journal of Accounting Information Systems
Benford's law and the limits of digit analysis
- Research Article
37
- 10.2308/jfar-51622
- Dec 18, 2015
- Journal of Forensic Accounting Research
False positives or “Type I errors,” wherein test results indicate fraud where none actually exists, have been described as a costly “cry wolf problem” in auditing practice. Benford's Law, which is used as one tool among many in screening for financial statement manipulation, is especially prone to false positives when applied to small and moderately sized datasets. Relying in part on Monte Carlo simulations, we describe with greater precision than extant literature the mathematical correlation between N and Mean Absolute Deviation (MAD), a statistic increasingly used for assessing deviation from Benford's Law. We recommend replacing MAD with an alternative, Excess MAD, which explicitly adjusts for N in estimating deviation from Benford's Law. Applying nonparametric, generalized additive modeling to public company financial statement numbers, we demonstrate the differing outcomes expected from Excess MAD and MAD and produce evidence suggesting that, despite Sarbanes-Oxley and Dodd-Frank legislation, Benford's Law conformity of public company financial statement numbers remained relatively stable across four decades beginning in 1970.
- Research Article
- 10.18267/j.aip.281
- Aug 11, 2025
- Acta Informatica Pragensia
Background: Benford's law is a statistical phenomenon that predicts the probability of a particular digit at a particular position in a number. This law has been successfully applied in a number of areas, such as accounting. In the area of scientometrics, research has been devoted mostly to journal data.Objective: This paper investigates the conformity of Benford's law with the citation counts of records retrieved from the Web of Science database. We evaluate the conformity levels with Benford's law in the complete dataset. We determine the effect of document type (article, proceedings paper and review), year of publication (2014-2018) and Web of Science categories (254 categories) on the level of conformity of the citation counts with Benford's law.Methods: The dataset of this research contains over 8.47 million records. All available records from the Web of Science were downloaded, so this set is the entire population of data available at the time of download. The distributions of the first significant digits in the citation counts of these records are compared with Benford's law. Mean absolute deviation (MAD) recommended by Nigrini (2012) and sum of squared deviations (SSD) recommended by Kossovsky (2015) are used to categorize the similarity of the citation counts to Benford's law.Results: The entire dataset of this study shows marginal conformity according to both MAD and SSD intervals (with a MAD value of 0.1257 and an SSD value of 29.9; a lower value indicates a better agreement). The review document type shows a high level of conformity, while proceedings paper shows a lower level. We found significant differences in conformity between Web of Science categories.Conclusion: This study mapped the level of conformity of the citation counts with Benford's law in data from the Web of Science database. Further directions for possible research are suggested.
- Research Article
3
- 10.17537/2022.17.230
- Nov 5, 2022
- Mathematical Biology and Bioinformatics
An empirical Benford's law which describes the probability of the appearance of certain first significant digits in many distributions taken from real life, is used to identify anomalies in various kinds of data. Our aim was to test Benford's law to assess the quality of mass preventive screening data on the example of bioelectrical impedance analysis (BIA) data from Moscow health centers. As was shown earlier, such a data is characterized by a high level of contamination by artificially generated and falsified data. A generated 2010–2019 database of BIA measurements contained 1361019 measurement records in the age range of the examined persons from 5 to 96 years. Application of the expert quality assessment algorithm, which was used as a reference for evaluation of the effectiveness of Benford analysis, revealed a high percentage of incorrect data (66.5 %) which was dominated by falsified data. To characterize the degree of the data compliance with Benford's law, the mean absolute deviations of the frequency distributions of the first and first two significant digits deviations from the proper values and chi-squared statistics for the tenth powers of the standardized resistance, reactance, and resistance index values were assessed for each health center. A significant correlation was observed between the data deviation from Benford's law and the percentage of incorrect data as provided by the expert quality assessment algorithm (ρmax = 0.66 and 0.62 for the mean absolute deviations and χ2 statistics, respectively, based on the resistance value and the first significant digit). It is suggested that deviation of the BIA data from Benford's law serves as a sufficient, but not a necessary, condition for their contamination. For those health centers, in which most of the incorrect data were represented by multiple measurements of the same person under the guise of different ones, the data were in good agreement with Benford's law. If the structure of incorrect data was dominated by measurements of the calibration block, software emulations of BIA measurements and outliers, then the use of Benford's law made it possible to effectively rank health centers by the level of data authenticity.
- Research Article
6
- 10.20473/jisebi.9.2.239-252
- Nov 1, 2023
- Journal of Information Systems Engineering and Business Intelligence
Background: Fraud in financial transaction is at the root of corruption issues recorded in organization. Detecting fraud practices has become increasingly complex and challenging. As a result, auditors require precise analytical tools for fraud detection. Grouping financial transaction data using K-Means Clustering algorithm can enhance the efficiency of applying Benford Law for optimal fraud detection. Objective: This study aimed to introduce Multiple Benford Law Model for the analysis of data to show potential concealed fraud in the audited organization financial transaction. The data was categorized into low, medium, and high transaction values using K-Means Clustering algorithm. Subsequently, it was reanalyzed through Multiple Benford Law Model in a specialized fraud analysis tool. Methods: In this study, the experimental procedures of Multiple Benford Law Model designed for public sector organizations were applied. The analysis of suspected fraud generated by the toolkit was compared with the actual conditions reported in audit report. The financial transaction dataset was prepared and grouped into three distinct clusters using the Euclidean distance equation. Data in these clusters was analyzed using Benford Law, comparing the frequency of the first digit’s occurrence to the expected frequency based on Benford Law. Significant deviations exceeding ±5% were considered potential areas for further scrutiny in audit. Furthermore, the analysis were validated by cross-referencing the result with the findings presented in the authorized audit organization report. Results: Multiple Benford Law Model developed was incorporated into an audit toolkit to automated calculations based on Benford Law. Furthermore, the datasets were categorized using K-Means Clustering algorithm into three clusters representing low, medium, and high-value transaction data. Results from the application of Benford Law showed a 40.00% potential for fraud detection. However, when using Multiple Benford Law Model and dividing the data into three clusters, fraud detection accuracy increased to 93.33%. The comparative results in audit report indicated a 75.00% consistency with the actual events or facts discovered. Conclusion: The use of Multiple Benford Law Model in audit toolkit substantially improved the accuracy of detecting potential fraud in financial transaction. Validation through audit report showed the conformity between the identified fraud practices and the detected financial transaction. Keywords: Fraud Detection, Benford’s Law, K-Means Clustering, Audit Toolkit, Fraudulent Practices.
- Research Article
3
- 10.46458/27121097.2018.24.37
- Dec 25, 2018
- Zbornik radova - Journal of Economy and Business
This paper presents the application of Benford's law in psychological pricing detection. Benford's law is naturally occurring law which states that digits have predictable frequencies of appearance with digit one having the highest frequency. Psychological pricing is one of the marketing pricing strategies directed on price setting which have the psychological impact on certain consumers. In order to investigate the application of Benford's law in psychological pricing detection, Benford's law is observed in the case of first and last digits. In order to inspect if the first and last digits of the observed prices are distributed according to the Benford’s law distribution or discrete uniform distribution respectively, mean absolute deviation measure, chi-square tests and Kolmogorov-Smirnov Z tests are used. Results of the analysis conducted on three price datasets have shown that the most dominating first digits are 1 and 2. On the other side, the most dominating last digits are 0, 5 and 9 respectively. The chi-square tests and Kolmogorov-Smirnov Z tests have showed that, at significance level of 5%, none of the three observed price datasets does have first digit distribution that fits to the Benford’s law distribution. Likewise, mean absolute deviation values have shown that there are large differences between the last digit distributions and the discrete uniform distribution implying psychological pricing in all price datasets.
- Research Article
1
- 10.46336/ijmsc.v3i2.205
- Apr 23, 2025
- International Journal of Mathematics, Statistics, and Computing
Bank fraud involves several actions such as manipulating duplication, forgery, changing accounting records and so on. This study aims to detect the potential for fraud in banking reports on customer final balances. The types of tests used to detect potential fraud in this study are the First Digit Test, Second Digit Test and First Two Digit Test Benford's Law. Benford's law states that the proportion of occurrences of numbers in certain numbers is not the same. In addition to the three Benford's Law tests, further statistical tests were carried out to determine the magnitude of the deviation between the actual proportion of occurrences and the expected proportion of Benford's Law using the Mean Absolute Deviation (MAD), Chi-square test, and Z test. This study uses secondary data on the final balance of customer deposits as of July 2023 as much as 20,105 data. The results showed that there were indications of fraud in the form of rounding and duplication of data on the customer's final balance. MAD results show that the proportion of occurrence of actual numbers is quite consistent with the proportion of occurrence of Benford's Law expectations. Based on the Z test, the balance that has the potential for fraud is the value with the first digit '5', the second digit '3' and the first two digits '23'. These numbers can be found in balances with a nominal value of Rp5000 and Rp5636 in the first digit '5', and Rp23038 in the second digit '3' and the first two digits '23'.
- Conference Article
- 10.53486/issc2025.22
- Jun 1, 2025
Benford's Law is a mathematical tool widely used in forensic accounting and auditing to detect numerical anomalies that may indicate fraud. This study explores its applicability in identifying financial fraud, with a focus on the MiMedx case, where deviations from the expected digit distribution revealed irregularities suggesting revenue manipulation. By applying Benford's Law, auditors identified an unusually high frequency of certain digits, raising suspicions of financial misrepresentation. The analysis confirms that statistical deviations from Benford's expected pattern can serve as a red flag for fraud detection. This method enables auditors and investigators to pinpoint suspect transactions efficiently, reducing the time and resources needed for fraud investigations. The case study demonstrates that integrating Benford's Law into financial auditing enhances fraud detection and strengthens financial transparency. Future research should explore its applications in emerging digital transactions and blockchain-based financial reporting.
- Conference Article
4
- 10.23919/cisti.2019.8760922
- Jun 1, 2019
The Benford's Law allows the identification of numbers manipulation indications and, therefore, it may be used as a technique to support the auditor in the identification of evidence of fraud. This study intends to evaluate the behavior of 27,058 Portuguese companies seeking to answer questions such as “Do the financial statement headings follow the Benford's Law?” and “Do companies with negative net results show more deviations from the distribution of the Benford's Law?”. The analysis showed that the turnover for the activity sectors under study (lodging and restoration and similars) does not always follows the Benford's Law and the behavior is different depending on whether the company presents positive or negative results. The analysis of the information considering different financial situations enables to reinforce the importance of applying the Benford's Law as a control tool to be used by auditors.
- Research Article
5
- 10.3414/me15-01-0076
- Jan 1, 2016
- Methods of Information in Medicine
Sophisticated anti-fraud systems for the healthcare sector have been built based on several statistical methods. Although existing methods have been developed to detect fraud in the healthcare sector, these algorithms consume considerable time and cost, and lack a theoretical basis to handle large-scale data. Based on mathematical theory, this study proposes a new approach to using Benford's Law in that we closely examined the individual-level data to identify specific fees for in-depth analysis. We extended the mathematical theory to demonstrate the manner in which large-scale data conform to Benford's Law. Then, we empirically tested its applicability using actual large-scale healthcare data from Korea's Health Insurance Review and Assessment (HIRA) National Patient Sample (NPS). For Benford's Law, we considered the mean absolute deviation (MAD) formula to test the large-scale data. We conducted our study on 32 diseases, comprising 25 representative diseases and 7 DRG-regulated diseases. We performed an empirical test on 25 diseases, showing the applicability of Benford's Law to large-scale data in the healthcare industry. For the seven DRG-regulated diseases, we examined the individual-level data to identify specific fees to carry out an in-depth analysis. Among the eight categories of medical costs, we considered the strength of certain irregularities based on the details of each DRG-regulated disease. Using the degree of abnormality, we propose priority action to be taken by government health departments and private insurance institutions to bring unnecessary medical expenses under control. However, when we detect deviations from Benford's Law, relatively high contamination ratios are required at conventional significance levels.
- Research Article
2
- 10.11591/ijphs.v13i1.23437
- Mar 1, 2024
- International Journal of Public Health Science (IJPHS)
Countries worldwide, including Indonesia, grappled with the unprecedented challenges brought about by the coronavirus disease (COVID-19) pandemic. Surveillance data vividly illustrates the profound effect of the COVID-19 pandemic in Indonesia. Both daily cases and deaths were raised, revealing the rapid transmission of the virus within communities. A quantitative study using a statistical approach was accomplished with secondary data to evaluate the quality of COVID-19 epidemiological surveillance data in Indonesia during the period between March 2020 to January 2021. The data was sourced from the World Health Organization (WHO) website using data reports on COVID-19 confirmed cases and deaths. A rapid tool called the first digit law or the fulfillment of Benford’s law was used to suggest good data quality for epidemiological surveillance. Data analysis used the Chi-squared test and the log-likelihood ratio test. Also, it displayed the difference in mean absolute deviation (MAD) to identify the proximity of the data and Benford’s law distribution. The results showed that both confirmed, and death case distributions were statistically non-conformity with Benford’s law distribution. In terms of quality data regarding the COVID-19 pandemic, the epidemiological surveillance system falls short of Benford’s law assumption. Benford's law has been acknowledged as an initial analysis that can expeditiously assess the performance of a surveillance system. The next phase of this study would be to conduct a complete evaluation suitably, especially in post-pandemic COVID-19.
- Research Article
7
- 10.5829/ije.2022.35.04a.02
- Jan 1, 2022
- International Journal of Engineering
Due to the ease of access to platforms that can be used by forgers to tamper digital documents, providing automatic tools for identifying forged images is now a hot research field in image processing. This paper presents a novel forgery detection algorithm based on variants of Benford's law. In the proposed method, Mean Absolute Deviation (MAD) feature is extracted using traditional Benford's law. Also, generalized Benford's law is used for mantissa distribution feature vector. In addition to Benford's law-based features, other statistical features are used to construct the final feature vector. Finally, support vector machine (SVM) with three different kernel functions is used to classify original and forged images. The method has been tested on two common image datasets (CASIA V1.0 and V2.0). The experimental results show that 0.27% and 0.21% improvements on CASIA V1.0 and CASIA V2.0 datasets were achieved, respectively in detection accuracy by the proposed method in comparison to best state-of-the-art methods. The proposed efficient algorithm has a simple implementation. Moreover, on the basis of Benford’s law rich features are extracted from images so that classification process is efficiently performed by a simple SVM classifier in a short time.
- Research Article
2
- 10.56943/joe.v2i3.367
- Aug 18, 2023
- Journal of Entrepreneurship
This research investigates Benford’s Law as a statistical instrument to detect financial fraud and errors. Benford's Law, also known as the First-Digit Law, states that lesser digits, specifically '1', frequently appear as the leading digit in numerous numerical datasets, and deviations from this distribution may indicate potential financial irregularities. This research examines literature demonstrating the application and efficacy of Benford’s Law in identifying numerical inconsistencies indicative of financial misconduct. This research investigates the use of Benford’s Law as a tool of detecting error and fraudulent activities within the sales data of two branches during 2022. A comprehensive dataset of 3098 records from Bandung and 539 records from Surabaya was collected after excluding certain data points that exhibited abnormalities. The application of Benford's statistical test discovered a discrepancy between the Benford probability and the observed probability, suggesting the presence of possible errors and frauds. The audit findings unveiled anomalies in pricing and instances of fraudulent activities in both locations, primarily due to pricing discrepancies and incorrect price inputs from sales orders. Furthermore, instances of fraud involved the manipulation of set prices for personal gain by salesmen. The results affirmed the hypothesis that a larger deviation between Benford’s probability and the observed probability corresponded with a higher incidence of error and fraud. However, it observes that Benford’s Law is not a stand-alone solution for detecting fraud, as not all financial datasets conform to it and deviations from the law only indicate the possibility of fraud, not confirm it. Therefore, the research suggests using Benford’s Law in conjunction with other data analysis and auditing techniques to conduct a comprehensive investigation. The conclusion of research emphasizes the significance of Benford’s Law in the field of forensic accounting and the need for multidimensional strategies for effective error and fraud detection.
- Research Article
1
- 10.37830/sjs.2023.1.03
- Jan 1, 2023
- Spanish Journal of Statistics
Most papers on Benford's Law primarily discuss either (1) the science and mathematics for explaining the law; or (2) how to apply the law, especially for detecting data manipulation and fraud; or (3) suggestions for statistical tests to determine if data conform to a Benford's distribution.Leonardo Campanelli's recent paper "Testing Benford's Law" strongly objects to a descriptive measure I discussed in my paper "The Promises and Pitfalls of Benford's Law"-as if that measure were intended for Benford's testing in the Category-3 sense relevant for Campanelli's paper (SJS, vol.4, 2022).This reflects a conflation of meanings for "testing" that is common in the Benford's literature, where many Category-2 papers claim they are applying (directly) conventional or new hypothesis tests as tools to detect fraud.Yet, fraud detection is a forensic and context-sensitive process, for which there is no set formula.In this paper, I clarify the sampling plan I had used earlier to collect and analyze a quasi-random sample of datasets, based on published criteria in the literature, to paint a tentative picture of how far real data vary, and in what ways, from abstract BL expectations.Further, I discuss simulations I have conducted to replicate and expand on my previous results.