Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Using simulation studies to evaluate statistical methods.

  • TL;DR
  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
TL;DR

Simulation studies involve generating data via pseudo-random sampling to evaluate statistical methods, enabling assessment of properties like bias. This tutorial offers structured guidance on design, analysis, and reporting, highlighting common issues and reviewing 100 recent articles to identify areas for improvement.

Abstract
Translate article icon Translate Article Star icon

Simulation studies are computer experiments that involve creating data by pseudo‐random sampling. A key strength of simulation studies is the ability to understand the behavior of statistical methods because some “truth” (usually some parameter/s of interest) is known from the process of generating the data. This allows us to consider properties of methods, such as bias. While widely used, simulation studies are often poorly designed, analyzed, and reported. This tutorial outlines the rationale for using simulation studies and offers guidance for design, execution, analysis, reporting, and presentation. In particular, this tutorial provides a structured approach for planning and reporting simulation studies, which involves defining aims, data‐generating mechanisms, estimands, methods, and performance measures (“ADEMP”); coherent terminology for simulation studies; guidance on coding simulation studies; a critical discussion of key performance measures and their estimation; guidance on structuring tabular and graphical presentation of results; and new graphical presentations. With a view to describing recent practice, we review 100 articles taken from Volume 34 of Statistics in Medicine, which included at least one simulation study and identify areas for improvement.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 7
  • 10.1371/journal.pone.0027974
The Communicability of Graphical Alternatives to Tabular Displays of Statistical Simulation Studies
  • Nov 23, 2011
  • PLoS ONE
  • Alex R Cook + 1 more

Simulation studies are often used to assess the frequency properties and optimality of statistical methods. They are typically reported in tables, which may contain hundreds of figures to be contrasted over multiple dimensions. To assess the degree to which these tables are fit for purpose, we performed a randomised cross-over experiment in which statisticians were asked to extract information from (i) such a table sourced from the literature and (ii) a graphical adaptation designed by the authors, and were timed and assessed for accuracy. We developed hierarchical models accounting for differences between individuals of different experience levels (under- and post-graduate), within experience levels, and between different table-graph pairs. In our experiment, information could be extracted quicker and, for less experienced participants, more accurately from graphical presentations than tabular displays. We also performed a literature review to assess the prevalence of hard-to-interpret design features in tables of simulation studies in three popular statistics journals, finding that many are presented innumerately. We recommend simulation studies be presented in graphical form.

  • Discussion
  • Cite Count Icon 3
  • 10.1016/j.jclinepi.2021.07.019
Prediction models: stepwise development and simultaneous validation is a step back
  • Aug 1, 2021
  • Journal of Clinical Epidemiology
  • Georg Heinze + 4 more

Prediction models: stepwise development and simultaneous validation is a step back

  • Research Article
  • 10.1093/ecco-jcc/jjad212.1056
P926 Improved standards of colonoscopy for Inflammatory Bowel Disease through implementation of key performance measures - A quality improvement initiative
  • Jan 24, 2024
  • Journal of Crohn's and Colitis
  • L Vaitiekunas + 4 more

Background Improving the quality of colonoscopy in inflammatory bowel disease is expected to improve clinical outcomes for patients. The European Society of Gastrointestinal Endoscopy (ESGE) recently published key performance measures for IBD colonoscopy assessment. This study aimed to assess the impact of implementing key performance measures and assess for sustainable improvement through a multimodal education intervention. Methods A baseline retrospective analysis was performed in patients with established IBD between June and August of 2022 at a tertiary hospital. The key performance measures included 1) pre-procedure metrics including indication, consent, and safety checklist (target of 100%) and 2) bowel preparation score, photo-documentation, disease activity scores, adequate biopsies, use of high-definition endoscopy and chromoendoscopy. There were six proceduralists involved and data was collected from electronic medical records including endoscopy and histopathology results. The ESGE performance measures were used to set minimum standards and we have adopted overall standards for our unit based on the ECCO and SCENIC consensus guidelines. Over 12 months, proceduralists and endoscopy nursing staff were engaged with educational interventions including didactic teaching for one hour and diagrammatic reminders in the endoscopy suites of the ESGE key performance measures. A post-implementation analysis was conducted at 1 year from baseline, from August to November of 2023. Results Baseline standards showed suboptimal performance in the use of disease activity scores and chromoendoscopy. Educational interventions were implemented and after 12 months, a repeat analysis of 50 consecutive patients showed significant improvement in all performance measures (see Table 1). Conclusion Quality metrics are important and underrecognised components in colonoscopy for IBD patients and form an integral part of improving patient care. Our study demonstrated that the implementation of the ESGE key performance measures is effective in improving the quality of colonoscopy assessment in IBD patients, identifying areas requiring further development and increasing the dysplasia detection rate. Acknowledging quality attrition over time, we recognise that regular teaching and education are required in addressing the challenge of sustainable long-term improvement.

  • Research Article
  • Cite Count Icon 3
  • 10.7939/r3bv7b69d
The effect of large ability differences on type I error and power rates using SIBTEST and TESTGRAF DIF detection procedures
  • Apr 1, 2002
  • University of Alberta Library
  • Andrea Gotzmann

A simulation study was conducted to examine the effect of large ability differences using two differential item functioning (DIF) detection procedures, SIBTEST and TESTGRAF. DIF items are hard to identify when group ability differences are large (Gotzmann, Vandenberghe, & Gierl, 2000; Hambleton & Rogers, 1989). This problem was investigated in the current study for the SIBTEST and TESTGRAF DIF detection procedures. Four ability differences (0.0, -1.0, -1.5, -2.0) and eight sample sizes (500/500, 750/1000, 1000/1000, 750/1500, 1000/1500, 1500/1500, 1000/2000, 2000/2000) were manipulated in a simulation study. Type I error and power rates were computed. The SIBTEST Type I error rates were inflated at the larger abilitjt differences. Conversely, the TESTGRAF Type I error rates remained low for most ability differences and sample sizes. The SIBTEST power rates remained high, even with larger ability differences. The TESTGRAF power rates dropped as ability differences were introduced. Ability Differences 3 The Effects of Large Ability Differences on Type I Error and Power Rates using the SIBTEST and TESTGRAF DIF Detection Procedures Educational practitioners and test developers often find large test scores differences when comparing examinees with diverse ethnic backgrounds (Berends & Koretz, 1996; Cameron, 1990; Freed le & Kostin, 1990; Scheuneman & Grima, 1997; Schmitt & Dorans, 1990). Reducing these differences is one goal in the educational reform movement (Barron & Koretz, 1996). These large test score differences are particularly noteworthy when Native and non-Native examinees are compared (Alberta Education, 1996; Gotzmann, Vandenberghe, & Gierl, 2000; Hambleton & Rogers, 1989; Vandenberghe & Gierl, 2001). Socioeconomic and cultural differences may contribute to these performance differences (Common & Frost, 1989; Hull, 1990; Trent & Gilman, 1985; Wood & Clay, 1996). However, few researchers have studied item-level outcomes which may explain why Native examinees score lower than non-Native examinees (Gotzmann et al., 2000; Hambleton & Rogers, 1989). Native examinee scores may be biased due to factors in test development. For example, Janzen (2000) and Krywaniuk and Das (1976) found that Native children are more likely to use simultaneous processing skills and non-Native children are more likely to use successive processing skills. If exams have a small number of items that illicit simultaneous processing skills, then these exams may put Native examinees at a disadvantage. Therefore, assessment of bias at the item level, and its contribution to the total test score differences, should be studied. Item bias can be estimated with different methods. Traditionally, item-level differences between groups have been assessed by comparing the proportion correct Ability Differences 4 for each group (Lord, 1980). However, this method has one major flaw. The proportion correct method compares all examinees, regardless of ability level. Thus, the proportion correct is dependent upon the sample of examinees (see Camilli & Shepard, 1994). To overcome this problem, statistical methods can be used to determine whether differential item functioning (DIF) is present. DIF occurs when examinees from different groups have a different probability of answering the ite-m,ebrrectly, after controlling for overall ability. In these comparisons, the majority group is called the reference group and the minority group is called the focal group. DIF methods are used to estimate bias by matching examinees on an internal measure of ability or overall test score performance and comparing these examinees at the item level. This approach removes total test score differences in the estimation process, which provides a stronger measure of the actual group differences on the item. There are many statistical procedures to estimate DIF including Item Response Theory (IRT) area measures (Lord, 1980; Thissen, Steinberg, & Wainer, 1988), MantelHaenszel (Holland & Thayer, 1988), Logistic Regression (Swaminathan & Rogers, 1990), Simultaneous Item Bias Test (SIBTEST; Shealy & Stout, 1993), and TESTGRAF (Ramsay, 1991, 2000). Most of these procedures have been used to identify DIF between ethnic groups. However, the SIBTEST and TESTGRAF procedures may be suitable when large ability differences are found. Further, both of these DIF detection procedures can be used with small sample sizes and both yield comparable DIF measures (Ramsay, 1991; 2000; Shealy & Stout, 1993). However, these procedures also have a noteworthy difference. SIBTEST uses a regression correction to estimate

  • Research Article
  • Cite Count Icon 36
  • 10.1097/ede.0b013e3182125cff
Designs for the Combination of Group- and Individual-level Data
  • May 1, 2011
  • Epidemiology
  • Sebastien Haneuse + 1 more

Studies of ecologic or aggregate data suffer from a broad range of biases when scientific interest lies with individual-level associations. To overcome these biases, epidemiologists can choose from a range of designs that combine these group-level data with individual-level data. The individual-level data provide information to identify, evaluate, and control bias, whereas the group-level data are often readily accessible and provide gains in efficiency and power. Within this context, the literature on developing models, particularly multilevel models, is well-established, but little work has been published to help researchers choose among competing designs and plan additional data collection. We review recently proposed "combined" group- and individual-level designs and methods that collect and analyze data at 2 levels of aggregation. These include aggregate data designs, hierarchical related regression, two-phase designs, and hybrid designs for ecologic inference. The various methods differ in (i) the data elements available at the group and individual levels and (ii) the statistical techniques used to combine the 2 data sources. Implementing these techniques requires care, and it may often be simpler to ignore the group-level data once the individual-level data are collected. A simulation study, based on birth-weight data from North Carolina, is used to illustrate the benefit of incorporating group-level information. Our focus is on settings where there are individual-level data to supplement readily accessible group-level data. In this context, no single design is ideal. Choosing which design to adopt depends primarily on the model of interest and the nature of the available group-level data.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 19
  • 10.1186/s41512-022-00124-y
A scoping methodological review of simulation studies comparing statistical and machine learning approaches to risk prediction for time-to-event data
  • Jun 2, 2022
  • Diagnostic and prognostic research
  • Hayley Smith + 3 more

BackgroundThere is substantial interest in the adaptation and application of so-called machine learning approaches to prognostic modelling of censored time-to-event data. These methods must be compared and evaluated against existing methods in a variety of scenarios to determine their predictive performance. A scoping review of how machine learning methods have been compared to traditional survival models is important to identify the comparisons that have been made and issues where they are lacking, biased towards one approach or misleading.MethodsWe conducted a scoping review of research articles published between 1 January 2000 and 2 December 2020 using PubMed. Eligible articles were those that used simulation studies to compare statistical and machine learning methods for risk prediction with a time-to-event outcome in a medical/healthcare setting. We focus on data-generating mechanisms (DGMs), the methods that have been compared, the estimands of the simulation studies, and the performance measures used to evaluate them.ResultsA total of ten articles were identified as eligible for the review. Six of the articles evaluated a method that was developed by the authors, four of which were machine learning methods, and the results almost always stated that this developed method’s performance was equivalent to or better than the other methods compared. Comparisons were often biased towards the novel approach, with the majority only comparing against a basic Cox proportional hazards model, and in scenarios where it is clear it would not perform well. In many of the articles reviewed, key information was unclear, such as the number of simulation repetitions and how performance measures were calculated.ConclusionIt is vital that method comparisons are unbiased and comprehensive, and this should be the goal even if realising it is difficult. Fully assessing how newly developed methods perform and how they compare to a variety of traditional statistical methods for prognostic modelling is imperative as these methods are already being applied in clinical contexts. Evaluations of the performance and usefulness of recently developed methods for risk prediction should be continued and reporting standards improved as these methods become increasingly popular.

  • Research Article
  • Cite Count Icon 12
  • 10.1057/palgrave.jors.2601651
Flexible versus fixed timetabling: a case study
  • Feb 1, 2004
  • Journal of the Operational Research Society
  • M Bazargan-Lari

This case study presents the timetabling problem of the Flight Training Department at Embry-Riddle Aeronautical University. The problem consists of scheduling the flight resources to students to various time blocks. This problem represents a well-studied field in operations research, mainly adopting variations of mathematical programming models. This paper initially presents the efforts towards developing a fixed timetable using optimization models for the case under study. It is, however, demonstrated that implementation of optimum solutions obtained using this approach cannot be sustained, mainly because of the dynamic nature of the governing parameters. A flexible and dynamic timetable utilizing the university computer network, allowing the instructors and students to make their own decentralized flexible timetables, is proposed. A simulation study is initiated to compare the performance measures under both timetables. The analysis shows that implementation of a flexible system generates higher utilization of flight resources as well as improving key performance measures.

  • Conference Article
  • 10.23919/cinc49843.2019.9005912
A New Graphical Method for Reporting Performance Results of a Diagnostic Test
  • Dec 30, 2019
  • Wang John

Reporting diagnostic performance results using standard performance measures, such as: sensitivity, specificity, and predictive values, have been a standard practice for decades. Issues with reporting using only numerical values include: 1) often only a subset of performance measures are reported, thus results could be misinterpreted and misused, 2) difficult to visualize the complex relationships of the reported performance measures. To overcome these shortcomings, a graphical presentation has been developed to further improve the test results reporting.The 2-dimentional performance graph uses line segment, area, ratio of line segments, ratio of areas, and sum of areas to represent most of the commonly used performance measures in this single graph.Advantages of the new graphic presentation are: 1) large number of performance measures can be presented and visualized simultaneously in a single graph, 2) allow the complex relationships of all performance measures to be understood more easily, 3) reduce the need to memorize some of the complex formulas for computing the performance measures, and 4) a great teaching tool in explaining the relationships of the commonly used performance measures.

  • Research Article
  • Cite Count Icon 29
  • 10.1111/caim.12370
Managing innovation performance: Results from an industry‐spanning explorative study on R&D key measures
  • May 7, 2020
  • Creativity and Innovation Management
  • Peter M Bican + 1 more

Research on R&D performance measures applied in firms is still scarce. Based on the established “R&D laboratory as a system” thinking, systematically derive and identify R&D department level key performance measures. Through a mixed‐method approach, grounded in (1) literature and (2) text analysis, 154 R&D performance measures were developed. Amongst those, an (3) online expert survey, as well as (4) three independent focus group workshops with >40 industry experts from more than ten industries identified and validated ten key R&D performance measures. All industry experts involved are members of an innovation network, additionally accounting for innovation network effects. In contrast to earlier research, some of the measures like degree of anticipation of internal customer needs were perceived both by the survey respondents and the focus groups as key measures, indicating that behavioral measures should not be excluded per se. However, the importance of external validity of R&D performance or indicators to measure performance in relation to activities outside the R&D department were not confirmed. Hence, we partly confirm the relevance of the original “R&D Lab” measures, contributing a more granular level, thereby drawing implications for future research and practice.

  • Research Article
  • Cite Count Icon 27
  • 10.1177/0037549711407781
A simulation study integrated with analytic hierarchy process (AHP) in an automotive manufacturing system
  • May 23, 2011
  • SIMULATION
  • Te Xu + 2 more

A variety of circumstances, such as developing a new product, changing the design of an existing product, changing the production volume, or changing the product mix, can drive the need for a manufacturing system redesign. Simulation technology has been widely used for evaluating manufacturing system design alternatives. Various performance measures including throughput, utilization of resources, lead-time, and work in process are obtained from simulation studies. However, the dimensions of performance measures are different, and it requires that the trade-offs between them should be considered. In addition, personal preferences in the selection of performance measures vary from person to person. Thus, the multi-criteria decision-making (MCDM) system is required to complement the simulation technology. This paper presents a case study that integrates a simulation study with analytic hierarchy process (AHP), applied to the design of a transmission case line in a Korean automotive factory. The simulation model is developed with QUEST®. Four performance measures (criteria) and seven design alternatives are considered utilizing AHP.

  • Research Article
  • Cite Count Icon 50
  • 10.1049/ip-epa:20030258
Moving block signalling dynamics: performance measures and re-starting queued electric trains
  • Jul 1, 2003
  • IEE Proceedings - Electric Power Applications
  • H Takeuchi + 2 more

From the functional point of view, signalling systems are generally evaluated on two basic criteria, namely their steady state and perturbed performance. The first of these essentially defines the capacity of a railway system, and the second relates to the cumulative disturbance experienced by all the trains on a system in response to a specific event, usually a delay imposed on a particular train. These two performance measures are investigated for both fixed-block signalling (FBS) and moving-block signalling (MBS) with other forms being included for comparative purposes. Steady state performance is analysed both mathematically and by using train movement simulators. The dynamic response to externally imposed disturbance is evaluated, using simulation studies. Some numerical measures of perturbed performance are proposed and the important trade-off between steady state and perturbed performance is discussed. A particular problem is the starting behaviour of a queue of electric trains when the leading train has stood still for a long time. This is the peak electrical demand problem caused by the characteristics of MBS. Peak demand reduction (PDR) techniques are thus proposed as solutions to the ‘starting problem’ and evaluated through simulation studies regarding peak reduction, delay and energy saving.

  • Research Article
  • Cite Count Icon 25
  • 10.1007/s00170-009-2424-x
A bi-criteria nonlinear fluctuation smoothing rule incorporating the SOM–FBPN remaining cycle time estimator for scheduling a wafer fab—a simulation study
  • Nov 20, 2009
  • The International Journal of Advanced Manufacturing Technology
  • Toly Chen + 1 more

This paper proposes a bi-criteria nonlinear fluctuation smoothing rule to further improve the performance of job scheduling in a wafer fabrication factory (wafer fab). The rule is based on the well-known fluctuation smoothing rules. First, the remaining cycle time of a job is estimated by applying the self-organization map–fuzzy back propagation network approach to improve the estimation accuracy. Second, two nonlinear forms of the fluctuation smoothing rules are obtained to enhance the balance and responsiveness. Third, the two nonlinear fluctuation smoothing rules are merged into a bi-criteria rule for considering two performance measures (average cycle time and cycle time variation) at the same time. Finally, the content of the bi-criteria rule can be tailored for the wafer fab and be scheduled with an adjustable factor. To evaluate the effectiveness of the proposed methodology, a production simulation was conducted. According to the experimental results, the proposed methodology outperformed some of the existing approaches by reducing the average cycle time and cycle time variation at the same time. In addition, the experimental results showed that the bi-criteria rule made it possible to improve one performance measure without raising the expense of another one.

  • Research Article
  • Cite Count Icon 48
  • 10.1016/j.cie.2006.08.007
A simulation study on the performance of pickup-dispatching rules for multiple-load AGVs
  • Sep 12, 2006
  • Computers & Industrial Engineering
  • Ying-Chin Ho + 1 more

A simulation study on the performance of pickup-dispatching rules for multiple-load AGVs

  • Research Article
  • Cite Count Icon 126
  • 10.1161/01.cir.0000435779.48007.5c
Synthesizing Lessons Learned From Get With The Guidelines
  • Oct 28, 2013
  • Circulation
  • A Gray Ellrodt + 11 more

The American Heart Association/American Stroke Association (AHA/ASA) is a trusted source of scientific information in cardiovascular medicine. The AHA/ASA has a longstanding commitment to support state-of-the-art scientific research in cardiovascular disease and stroke. The AHA/ASA has also developed a leadership role in translating cardiovascular science into internationally respected guidelines. In 2000, however, the AHA/ASA concluded that, to provide maximal benefit for patients with cardiovascular disease and those at risk, it needed to develop a rigorous approach to translating its guidelines into clinical practice. The result was a comprehensive suite of programs collectively called Get With The Guidelines (GWTG). Modeled in part on the University of California, Los Angeles Cardiovascular Hospitalization Atherosclerosis Management Program (CHAMP),1,2 GWTG was successfully piloted by the AHA in Massachusetts.3 Based on the success of the Massachusetts pilot, the AHA committed significant human and financial resources to extend the program across the United States. This commitment included the development of a national steering committee composed of AHA/ASA volunteers, and the addition of multiple modules including hospital-based management of coronary artery disease, heart failure, stroke, and resuscitation after in-hospital cardiac arrest. The scientific foundation of the program is the best evidence from the latest American College of Cardiology/AHA/ASA guidelines. GWTG staff work with participating hospitals to implement these guidelines by using AHA/ASA quality improvement professional consultation, workshops, and Webinars. In addition, the AHA developed sophisticated clinical databases (registries) through which hospitals and physicians collect information in real time for the assessment of quality, regional, and national benchmarking, national recognition, and the generation of new science. The AHA underwrites a portion of the costs associated with the technology platform and data collection tools to reduce the financial burden on participating sites. From the 4 GWTG disease-specific registries, >200 articles have been published in peer-reviewed journals. …

  • Research Article
  • Cite Count Icon 1
  • 10.1080/09720510.2020.1862961
Evaluation of statistical methods for assessing reliability from degradation data : A simulation study
  • Jun 8, 2021
  • Journal of Statistics and Management Systems
  • Herbert Hove

This paper deals with the evaluation of statistical methods for assessing product reliability from degradation data. The evaluation is based on a simulation study. Benefits of assessing reliability from degradation data over censored failure time data are demonstrated using data from the first simulation run. Bias, precision and coverage probability calculated from the 105 simulated samples are the performance measures used in the evaluation. Analysis of a real gallium arsenide (GaAs) laser data set for telecommunication systems emphasizes the practical advantages of assessing reliability from degradation data.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant