Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Propensity Score Matching in Accounting Research

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

ABSTRACT Propensity score matching (PSM) has become a popular technique for estimating average treatment effects (ATEs) in accounting research. In this study, we discuss the usefulness and limitations of PSM relative to more traditional multiple regression (MR) analysis. We discuss several PSM design choices and review the use of PSM in 86 articles in leading accounting journals from 2008–2014. We document a significant increase in the use of PSM from zero studies in 2008 to 26 studies in 2014. However, studies often oversell the capabilities of PSM, fail to disclose important design choices, and/or implement PSM in a theoretically inconsistent manner. We then empirically illustrate complications associated with PSM in three accounting research settings. We first demonstrate that when the treatment is not binary, PSM tends to confine analyses to a subsample of observations where the effect size is likely to be smallest. We also show that seemingly innocuous design choices greatly influence sample composition and estimates of the ATE. We conclude with suggestions for future research considering the use of matching methods. Data Availability: All data used are available from sources cited in the text.

Similar Papers
  • Front Matter
  • Cite Count Icon 17
  • 10.1016/j.jtcvs.2020.10.120
Randomized trials, observational studies, and the illusive search for the source of truth
  • Nov 10, 2020
  • The Journal of Thoracic and Cardiovascular Surgery
  • Mario Gaudino + 3 more

Randomized trials, observational studies, and the illusive search for the source of truth

  • Abstract
  • 10.1016/j.psyneuen.2015.07.490
Examining the nature and directionality of the relationship between children's anxiety symptoms and the immunoendocrine systems
  • Aug 8, 2015
  • Psychoneuroendocrinology
  • Denise Ma + 2 more

Examining the nature and directionality of the relationship between children's anxiety symptoms and the immunoendocrine systems

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 13
  • 10.1038/s41598-022-13344-5
Correlation between air pollution and prevalence of conjunctivitis in South Korea using analysis of public big data
  • Jun 16, 2022
  • Scientific Reports
  • Sanghyu Nam + 6 more

This study investigated how changes in weather factors affect the prevalence of conjunctivitis using public big data in South Korea. A total of 1,428 public big data entries from January 2013 to December 2019 were collected. Disease data and basic climate/air pollutant concentration records were collected from nationally provided big data. Meteorological factors affecting eye diseases were identified using multiple linear regression and machine learning analysis methods such as extreme gradient boosting (XGBoost), decision tree, and random forest. The prediction model with the best performance was XGBoost (1.180), followed by multiple regression (1.195), random forest (1.206), and decision tree (1.544) when using root mean square error (RMSE) values. With the XGBoost model, province was the most important variable (0.352), followed by month (0.289) and carbon monoxide exposure (0.133). Other air pollutants including sulfur dioxide, PM10, nitrogen dioxides, and ozone showed low associations with conjunctivitis. We identified factors associated with conjunctivitis using traditional multiple regression analysis and machine learning techniques. Regional factors were important for the prevalence of conjunctivitis as well as the atmosphere and air quality factors.

  • Front Matter
  • Cite Count Icon 54
  • 10.1016/j.ijrobp.2020.02.638
Understanding Propensity Score Analyses
  • Jun 9, 2020
  • International Journal of Radiation Oncology*Biology*Physics
  • Nafisha Lalani + 2 more

Understanding Propensity Score Analyses

  • Research Article
  • Cite Count Icon 83
  • 10.1002/bimj.201600094
Propensity scores based methods for estimating average treatment effect and average treatment effect among treated: A comparative study.
  • Apr 24, 2017
  • Biometrical Journal
  • Younathan Abdia + 4 more

Propensity score based statistical methods, such as matching, regression, stratification, inverse probability weighting (IPW), and doubly robust (DR) estimating equations, have become popular in estimating average treatment effect (ATE) and average treatment effect among treated (ATT) in observational studies. Propensity score is the conditional probability receiving a treatment assignment with given covariates, and propensity score is usually estimated by logistic regression. However, a misspecification of the propensity score model may result in biased estimates for ATT and ATE. As an alternative, the generalized boosting method (GBM) has been proposed to estimate the propensity score. GBM uses regression trees as weak predictors and captures nonlinear and interactive effects of the covariate. For GBM-based propensity score, only IPW methods have been investigated in the literature. In this article, we provide a comparative study of the commonly used propensity score based methods for estimating ATT and ATE, and examine their performances when propensity score is estimated by logistic regression and GBM, respectively. Extensive simulation results indicate that the estimators for ATE and ATT may vary greatly due to different methods. We concluded that (i) regression may not be suitable for estimating ATE and ATT regardless of the estimation method of propensity score; (ii) IPW and stratification usually provide reliable estimates of ATT when propensity score model is correctly specified; (iii) the estimators of ATE based on stratification, IPW, and DR are close to the underlying true value of ATE when propensity score is correctly specified by logistic regression or estimated using GBM.

  • Research Article
  • Cite Count Icon 7
  • 10.1007/s11250-020-02526-w
Application of principal component analysis for predicting body weight of Ethiopian indigenous chicken populations.
  • Jan 8, 2021
  • Tropical Animal Health and Production
  • Fikrineh Negash

Under smallholder management conditions, where weighing scale is not readily available, body weight (BW) can be predicted from morphometric measurements using multiple regression. However, the statistical interpretation of the regression parameters estimated by least squares is difficult due to multicollinearity problems. Thus, principal component analysis (PCA) was used to predict BW from body measurements, including body length (BL), chest circumference (CC), shank length (SL), and shank circumference (SC) of Ethiopian indigenous chicken populations reared by smallholder farmers. The effectiveness of this technique was also compared with the traditional multiple regression analysis (MRA). Measurements were taken from 134 male and 487 female chickens. The whole dataset was partitioned into two portions, namely, training and testing datasets, for model comparison and validation purposes. The training dataset, which consisted of 75% of the dataset, was used to develop the model, and the testing dataset, which consisted of 25% of the dataset, was used to validate the model. The PCA results revealed that the variables for body measurements were represented by PC1 and PC2 in male birds and PC1, PC2, and PC3 in female birds. Regression models developed using scores derived from these PCs explained 88 and 69% of total variation in BW in male and female birds, respectively. Compared with traditional MRA, regression models generated using the PCA procedure were more accurate in predicting BW. Thus, the results of the present study could not only be used for predicting BW of Ethiopian indigenous chickens but also in their genetic selection programs.

  • Discussion
  • Cite Count Icon 7
  • 10.1097/ede.0000000000000485
Commentary: Can a Quasi-experimental Design Be a Better Idea than an Experimental One?
  • Jul 1, 2016
  • Epidemiology
  • Jeremy Alexander Labrecque + 1 more

Commentary: Can a Quasi-experimental Design Be a Better Idea than an Experimental One?

  • Research Article
  • Cite Count Icon 68
  • 10.1177/0962280213507034
Propensity score estimators for the average treatment effect and the average treatment effect on the treated may yield very different estimates
  • Sep 30, 2016
  • Statistical Methods in Medical Research
  • R Pirracchio + 5 more

Propensity score matching is typically used to estimate the average treatment effect for the treated while inverse probability of treatment weighting aims at estimating the population average treatment effect. We illustrate how different estimands can result in very different conclusions. We applied the two propensity score methods to assess the effect of continuous positive airway pressure on mortality in patients hospitalized for acute heart failure. We used Monte Carlo simulations to investigate the important differences in the two estimates. Continuous positive airway pressure application increased hospital mortality overall, but no continuous positive airway pressure effect was found on the treated. Potential reasons were (1) violation of the positivity assumption; (2) treatment effect was not uniform across the distribution of the propensity score. From simulations, we concluded that positivity bias was of limited magnitude and did not explain the large differences in the point estimates. However, when treatment effect varies according to the propensity score (E[Y(1)-Y(0)|g(X)] is not constant, Y being the outcome and g(X) the propensity score), propensity score matching ATT estimate could strongly differ from the inverse probability of treatment weighting-average treatment effect estimate. We show that this empirical result is supported by theory. Although both approaches are recommended as valid methods for causal inference, propensity score-matching for ATT and inverse probability of treatment weighting for average treatment effect yield substantially different estimates of treatment effect. The choice of the estimand should drive the choice of the method.

  • Research Article
  • Cite Count Icon 158
  • 10.1016/j.jretconser.2019.101929
FsQCA versus regression: The context of customer engagement
  • Sep 2, 2019
  • Journal of Retailing and Consumer Services
  • David Gligor + 1 more

FsQCA versus regression: The context of customer engagement

  • Research Article
  • Cite Count Icon 37
  • 10.1016/j.measurement.2019.106975
Predicting deformational properties of Indian coal: Soft computing and regression analysis approach
  • Aug 23, 2019
  • Measurement
  • Debanjan Guha Roy + 1 more

Predicting deformational properties of Indian coal: Soft computing and regression analysis approach

  • Research Article
  • 10.3760/cma.j.issn.1673-4181.2017.04.003
An adaptive algorithm for single-trial P300 evoked potential extraction based on dynamic feature library
  • Aug 28, 2017
  • International Journal of Biomedical Engineering
  • Hanlei Li + 5 more

Objective The single-trial extraction method of evoked potential has been one of the problems in EEG information processing field. According to the characteristics of somatosensory evoked electroencephalogram (EEG) with low signal-to-noise ratio and large parameter variation between trials, a novel single-trial extraction method for evoked potentials was proposed. This method aims to further improve the accuracy and characteristics of the single-trial extraction algorithm, preserve more dynamic characteristics between trials, and improve the estimation accuracy. Methods Based on wavelet filtering and multiple linear analysis, a new single-trial extraction method for EEG P300 parameters was proposed by applying the adaptive dynamic feature library. Four groups of wavelet filtered evoked EEG data were randomly selected, and used to build the feature library using overlapping average method and principal component analysis. For the single-trial extracted EEG data, the component with the highest correlation coefficient related with the current data was selected as the independent variable from the feature library, and the relevant multiple linear regression analysis was conducted. The single-trial evoked potential signal was reconstructed by the regression analysis results, in which the key features such as latency and amplitude were automatically extracted. Results Compared with the benchmark values determined by experts, the proposed algorithm can obtain more accurate estimation values of latency and amplitude in P300 components. The average difference of latency and amplitude by the proposed algorithm is (11.16±8.60) ms and (1.40±1.34) μV, respectively. These two values obtained by the proposed algorithm are much closer to that obtained by the commonly used overlapping average method of (23.26 ± 25.76) ms and (2.52 ± 2.50) μV, respectively. These results show that the proposed algorithm has significant advantages comparing with the traditional multiple linear regression analysis algorithm. Conclusions The dynamic updating principal component sample library of EEG data was applied to wavelet filtering and multiple linear regression, thus the dynamic characteristics were effectively preserved, and the accuracy of parameter estimation was improved. Key words: Multiple linear regression with dispersion terms; Evoked electroencephalogram; P300; Principal component analysis

  • Research Article
  • 10.1161/str.52.suppl_1.p106
Abstract P106: Population-Based Drug Repurposing to Treat Thrombosis in COVID-19 Patients
  • Mar 1, 2021
  • Stroke
  • Yejin Kim + 3 more

Introduction: Patients with coronavirus disease 2019 (COVID-19) have an increased risk of thrombosis. Our objective is to obtain population-level treatment effects of drugs on treating thrombosis in COVID-19. Methods: We conducted a retrospective analysis of Optum electronic health records (EHRs) with 34,043 hospitalized COVID-19 patients. We identified case-patient with thrombosis (stroke, deep vein thrombosis, pulmonary embolism, and myocardial infarction) using PheWas codes. The propensity score matching was used to select comparable control patients who survived without any thrombosis based on demographics and admission status (temperature and SpO 2 level). We computed the average treatment effect (ATT) for medication using advanced inverse propensity score weighting based on pre-treatment conditions (i.e., comorbidities in the last 6 months and medications in the last 2 months before hospitalization). Results: We identified 2,446 case-patients with thrombosis and 5,020 comparable control patients. There were a total of 540 drugs that were administered in at least 80 patients. We calculated the 540 drugs’ ATT coefficient. As a result, 23 drugs had a positive ATT coefficient with a p -value of less than 0.05. After filtering out commonly prescribed symptomatic drugs (e.g., Acetaminophen, Guaifenesin, and Ondansetron), we highlight the following drugs with statistically significant treatment effects: Atorvastatin (ATT=0.34), Ceftriaxone (ATT=0.26), Levothyroxine (ATT=0.26), Albuterol (ATT=0.25), Azithromycin (ATT=0.23), Enoxaparin (ATT=0.20), and Metformin (ATT=0.20). Conclusions: In this preliminary work, we identified anti-thrombotic drugs (Enoxaparin) but also anti-inflammatory drugs (Atorvastatin, Metformin) and possibly antibiotics that have a significant treatment effect in COVID-19 patients that could reduce risk of thrombosis. We also observed that several anti-thrombotic drugs (Apixaban and Ticagrelor) had negative treatment effects, which was partly due to an imbalance in pre-treatment conditions. Our future work is to incorporate more extensive data (such as lab tests and vital signs) into the propensity scores to better capture the severity of admission status.

  • Research Article
  • Cite Count Icon 7
  • 10.2139/ssrn.3272396
Estimating Average Treatment Effects With Propensity Scores Estimated With Four Machine Learning Procedures: Simulation Results in High Dimensional Settings and With Time to Event Outcomes
  • Nov 16, 2018
  • SSRN Electronic Journal
  • Kip Brown + 2 more

Estimating Average Treatment Effects With Propensity Scores Estimated With Four Machine Learning Procedures: Simulation Results in High Dimensional Settings and With Time to Event Outcomes

  • Discussion
  • 10.1097/aln.0000000000001786
In Reply.
  • Oct 1, 2017
  • Anesthesiology
  • Daniel I Mcisaac + 2 more

We thank Drs. Hwang and Jeon and Drs. Kehlet and Jørgensen for their letters and welcome the opportunity to discuss the strengths and limitations of our study.1As stated in the letter from Drs. Hwang and Jeon and acknowledged in our article,1 we were unable to identify whether each nerve block studied was actually clinically effective. When considered from the perspective of an explanatory research question, this is clearly a limitation. However, because the aim of our study was comparative effectiveness, our specific objective was in the realm of pragmatic research, that is, how effective and generalizable might the intervention be in real-world practice.2 From this perspective, we hope that our measures of association provide useful insights into the impact that the peripheral nerve blocks have on system outcomes across a generalizable large sample of patients across an entire healthcare system.With respect to the assertion by Drs. Hwang and Jeon that our lack of control for intraoperative and postoperative variables and complications is a limitation, we would argue the contrary. In observational comparative effectiveness research, efforts must be made to adjust for indication bias and confounding bias (among other sources). When selecting variables that may be confounders, one must ensure that they meet the definition of a confounder, specifically that they differentially impact exposure (i.e., receipt of a block), differentially impact outcome, and are not on the causal pathway.3 Therefore, although complications may contribute to differences in length of stay (LOS), they are not true confounders because they occur after exposure and are likely on the causal pathway to prolonged LOS. Furthermore, it has been shown that control for variables such as these that are not true confounders can lead to spurious associations.4Finally, we agree with Drs. Hwang and Jeon that the choice of analytic approach when performing propensity score–based analyses impacts interpretation of study results.5 Specifically, matched analyses such as ours estimate the average treatment effect in the treated (ATT), because some individuals are excluded if they received treatment but no adequate match was available or if they were untreated and again went unmatched to a treated subject. Although this may decrease overall generalizability, it may also decrease bias. In contrast, methods such as inverse probability of treatment weighting (IPTW) or regression analysis provide an average treatment effect (ATE), that is, what might happen if the entire population was shifted from untreated to treated.6 Although the ATT and ATE are typically similar in direction and magnitude, this is not always the case. In fact, in the case of IPTW, including individuals who were treated despite a very low propensity for treatment can greatly over-weight their contribution to the analysis, especially if extreme tails of the distribution are not trimmed.5 Furthermore, matched analyses can provide an estimate of the absolute risk difference, as opposed to IPTW and regression-based approaches that are typically limited to estimating relative outcome differences. Lastly, in our sensitivity analysis we used a regression-based multilevel multivariable regression analysis, which estimated an ATE for single shot blocks that was identical in direction and magnitude to the ATT estimated from the propensity score–matched analysis.We would also like to thank Drs. Kehlet and Jørgensen for their commentary regarding our publication1 and in particular their interest in promoting improvements in reporting, analysis, and overall research efforts related to LOS. First, we agree that different patterns of care between jurisdictions or individual hospitals can skew LOS findings. As Hart et al.7 outlined in an analysis of Canadian versus American total joint arthroplasty outcomes, LOS in Canadian hospitals tends to be approximately 1.3 to 1.4 days longer, a finding that may be attributable to a 21 to 27% increase in rates of discharge to short-term rehabilitation from American hospitals. Data from Hart et al.7 also suggest a mean LOS after joint replacement in Canadian hospitals of slightly more than four days, a figure consistent with mean LOS reported in our study, which included a larger cross-section of hospitals than would have been included in the National Surgical Quality Improvement Program data file.Regarding differences in practice between hospitals, we fully acknowledge that our data sources do not allow us to measure whether specific fast-track processes of care were used at certain hospitals and for specific patients in our study; this is a limitation. For this reason, we ensured that all of our analyses accounted for clustering of patients within hospitals to allow us to account for unmeasured variation between hospitals, both in the use of perioperative processes of care as well as discharge patterns. In our propensity score–matched analysis this involved direct matching within hospitals along with a propensity score, a method that has been shown to decrease both bias and error in estimating causal effects relative to simply matching on the propensity score.8 In our sensitivity analysis, which used regression analysis, we accounted for clustering of patients in hospitals using a multivariable-adjusted generalized linear model and generalized estimating equation methods. We certainly encourage the use of analytic strategies that account for hierarchal data in all comparative effectiveness research where between-center variation is a consideration.In summary, across a universal healthcare system we report the population-based association between peripheral nerve block exposure and LOS using best-practice methods for comparative effectiveness research and report a LOS consistent with other reports from our jurisdiction. We agree that our data, like any observational data set, have limitations that must be considered when appraising our findings. We also agree that understanding why patients remain in the hospital after surgery is a high-priority area of research and that minimizing variation and instituting best practices should lead to improved patient and system outcomes.The authors declare no competing interests.

  • Research Article
  • 10.1002/hpm.3938
Heterogeneous Effects of Decreasing the Cost-Sharing for Outpatient Care on Health Outcomes in China: A Propensity Score Matching and Causal Machine Learning Approach.
  • May 4, 2025
  • The International journal of health planning and management
  • Tao Zhang + 3 more

To improve accessibility and financial support for outpatient services, China introduced a scheme to decrease cost-sharing for outpatient care under the Urban Employee Basic Medical Insurance. This study evaluates the health impacts of this policy and examines its heterogeneous effects. Utilising data from the 2018 China Health and Retirement Longitudinal Study, we analysed 2896 individual-level observations across 105 prefectures. Propensity score matching and a causal forest model were applied to evaluate the effects on chronic disease status, body pain, self-rated health, and hospitalisation, while accounting for various demographic, socioeconomic, residential, health-related behaviours, and prefecture-specific factors. The reduction in cost-sharing was significantly linked to decreased probabilities of chronic disease (Average Treatment Effect (ATE)=-0.0619, p<0.01), body pain (ATE=-0.0715, p<0.05), and hospitalisation (ATE=-0.0592, p<0.001), as well as improved self-rated health (ATE=0.1557, p<0.001). These benefits may be attributed to reduced out-of-pocket payments for outpatient care (ATE=-287.6112, p<0.01) and increased outpatient visits (ATE=0.0414 visits, p<0.05). Causal forest analyses revealed that older individuals, those with higher educational attainment, higher household income, urban residents, and those engaging in healthier behaviours exhibited larger treatment effects. Decreasing outpatient cost-sharing in China has beneficial health outcomes, with variations in its impact based on socio-economic status and health behaviours. It is advisable to further increase reimbursement rates and broaden benefit packages for outpatient care, while addressing the unequal distribution of benefits.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant