Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Sample Size Guidelines for Logistic Regression from Observational Studies with Large Population: Emphasis on the Accuracy Between Statistics and Parameters Based on Real Life Clinical Data

  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

BackgroundDifferent study designs and population size may require different sample size for logistic regression. This study aims to propose sample size guidelines for logistic regression based on observational studies with large population.MethodsWe estimated the minimum sample size required based on evaluation from real clinical data to evaluate the accuracy between statistics derived and the actual parameters. Nagelkerke r-squared and coefficients derived were compared with their respective parameters.ResultsWith a minimum sample size of 500, results showed that the differences between the sample estimates and the population was sufficiently small. Based on an audit from a medium size of population, the differences were within ± 0.5 for coefficients and ± 0.02 for Nagelkerke r-squared. Meanwhile for large population, the differences are within ± 1.0 for coefficients and ± 0.02 for Nagelkerke r-squared.ConclusionsFor observational studies with large population size that involve logistic regression in the analysis, taking a minimum sample size of 500 is necessary to derive the statistics that represent the parameters. The other recommended rules of thumb are EPV of 50 and formula; n = 100 + 50i where i refers to number of independent variables in the final model.

Similar Papers
  • Research Article
  • Cite Count Icon 71
  • 10.1111/j.1365-2044.2012.07155.x
If it hasn’t failed, does it work? On ‘the worst we can expect’ from observational trial results, with reference to airway management devices
  • May 7, 2012
  • Anaesthesia
  • J J Pandit

If it hasn’t failed, does it work? On ‘the worst we can expect’ from observational trial results, with reference to airway management devices

  • Research Article
  • Cite Count Icon 301
  • 10.1177/0013164407310131
Sample Sizes When Using Multiple Linear Regression for Prediction
  • Nov 9, 2007
  • Educational and Psychological Measurement
  • Gregory T Knofczynski + 1 more

When using multiple regression for prediction purposes, the issue of minimum required sample size often needs to be addressed. Using a Monte Carlo simulation, models with varying numbers of independent variables were examined and minimum sample sizes were determined for multiple scenarios at each number of independent variables. The scenarios arrive from varying the levels of correlations between the criterion variable and predictor variables as well as among predictor variables. Two minimum sample sizes were determined for each scenario, a good and an excellent prediction level. The relationship between the squared multiple correlation coefficients and minimum necessary sample sizes were examined. A definite relationship, similar to a negative exponential relationship, was found between the squared multiple correlation coefficient and the minimum sample size. As the squared multiple correlation coefficient decreased, the sample size increased at an increasing rate. This study provides guidelines for sample size needed for accurate predictions.

  • Research Article
  • Cite Count Icon 3
  • 10.1016/j.ijpara.2025.05.003
The sample size matters: evaluating minimum and reasonable values in prevalence studies.
  • Nov 1, 2025
  • International journal for parasitology
  • Volodimir Sarabeev + 5 more

The sample size matters: evaluating minimum and reasonable values in prevalence studies.

  • Research Article
  • Cite Count Icon 12
  • 10.1097/01.hj.0000396585.52118.6b
Enough is enough: A primer on power analysis in study designs
  • Apr 1, 2011
  • The Hearing Journal
  • Chi-Chuen Lau + 1 more

Enough is enough: A primer on power analysis in study designs

  • Discussion
  • Cite Count Icon 6
When calculation of minimum sample size is not justified
  • Mar 1, 2011
  • Hepatitis Monthly
  • Mohamad Amin Pourhoseingholi + 1 more

Dear Editor, We read with interest the paper of Estevez, et al, recently published in Hepatitis Monthly [1]. That article was a Letter to the Editor on the study conducted by Nikui Nejad et al [2]. One of the comments made by Estevez et al was on the adequacy of the size of the studied sample. They believed that the sample size was not adequate based on the provided data and the type of study [1]. In the reply from the authors, a familiar mathematical equation which is frequently used for study designs was presented to explain the way the sample size was calculated in Nikui Nejad's study [1]. In spite of doubts in the statistical analyses used in their study, we should emphasize that the study of Nikui Nejad, et al, was clearly a clinical trial involving comparison between two vaccines. In that study, the sample size did not depend on the of vaccine response rate in the population-instead, it did depend on the difference between the response rates from the two vaccines. The authors also used an incorrect assumption for the calculation of minimum sample size according the prevalence approach and assumed a large d with a small P! We know that in clinical trial, if the sample size is too small, a well-conducted study may fail to answer its research hypothesis or to detect important effects and associations. Therefore, correct assumptions for the calculation of minimum sample size are of paramount importance in conducting research. Approaches for estimating sample size and performing power analysis depend primarily on the study design and the main outcome measure [3]. Clinical trials should be large enough to detect reliably the smallest possible differences in the primary outcomes. It is not uncommon for studies to be underpowered, failing to detect even large treatment effects because of inadequate sample size [4]. The minimum information needed to calculate sample size for a randomized controlled trial includes the study power (1-s), the level of significance (α), the underlying event rate in the population and the effect size. The calculated minimum sample size should then be adjusted for other factors, including expected compliance rates and, less commonly, an unequal allocation ratio. The objectives and outcome measures of the study must be clearly stated, and the information used in calculating the minimum sample size should reflect as closely as possible the type of data that will be gathered from the proposed trial [5]. Sample size calculation is an important part of any clinical trials and a professional statistician is the best person to be asked for help at the time of planning a research project. However, researchers must be prepared to provide the necessary information so that the sample size can be determined [6]. There are many statistical books on the methods for sample size calculation in medical studies [7]. There are also several software programs available to help with sample size calculations. While these programs are easy to use, investigators should consult biostatisticians at the design stages of their projects and any article containing even the most elementary statistical procedure should be reviewed by an expert biostatistician.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/cibcb.2011.5948461
Derivation of minimum best sample size from microarray data sets: A Monte Carlo approach
  • Apr 1, 2011
  • Chengpeng Bi + 2 more

NCBI has been accumulating a large repository of microarray data sets, namely Gene Expression Omnibus (GEO). GEO is a great resource enabling one to pursue various biological and pathological questions. The question we ask here is: given a set of gene signatures and a classifier, what is the best minimum sample size in a clinical microarray research that can effectively distinguish different types of patient responses to a therapeutic drug. It is difficult to answer the question since the sample size for most microarray experiments stored in GEO is very limited. This paper presents a Monte Carlo approach to simulating the best minimum microarray sample size based on the available data sets. Support Vector Machine (SVM) is used as a classifier to compute prediction accuracy for different sample size. Then, a logistic function is applied to fit the relationship between sample size and accuracy whereby a theoretic minimum sample size can be derived.

  • Research Article
  • Cite Count Icon 5
  • 10.4300/jgme-d-15-00512.1
Questions Program Directors Need to Answer Before Using Resident Clinical Performance Data.
  • Oct 1, 2016
  • Journal of Graduate Medical Education
  • Sarah Gebauer + 1 more

Questions Program Directors Need to Answer Before Using Resident Clinical Performance Data.

  • Research Article
  • Cite Count Icon 1576
  • 10.1207/s15327574ijt0502_4
Minimum Sample Size Recommendations for Conducting Factor Analyses
  • Jun 1, 2005
  • International Journal of Testing
  • Daniel J Mundfrom + 2 more

There is no shortage of recommendations regarding the appropriate sample size to use when conducting a factor analysis. Suggested minimums for sample size include from 3 to 20 times the number of variables and absolute ranges from 100 to over 1,000. For the most part, there is little empirical evidence to support these recommendations. This simulation study addressed minimum sample size requirements for 180 different population conditions that varied in the number of factors, the number of variables per factor, and the level of communality. Congruence coefficients were calculated to assess the agreement between population solutions and sample solutions generated from the various population conditions. Although absolute minimums are not presented, it was found that, in general, minimum sample sizes appear to be smaller for higher levels of communality; minimum sample sizes appear to be smaller for higher ratios of the number of variables to the number of factors; and when the variables-to-factors ratio exceeds 6, the minimum sample size begins to stabilize regardless of the number of factors or the level of communality.

  • Research Article
  • Cite Count Icon 25
  • 10.1080/19439962.2014.978963
Minimum sample sizes for estimating reliable Highway Safety Manual (HSM) calibration factors
  • Oct 31, 2014
  • Journal of Transportation Safety & Security
  • Priyanka Alluri + 2 more

ABSTRACTThe Highway Safety Manual (HSM) assists state and local agencies in improving highway safety by moving toward statistically proven quantitative analyses. The manual recommends using Empirical Bayes (EB) method with locally derived calibration factors to predict the agency's safety performance. It recommends deriving calibration factors using randomly selected 30 to 50 roadway sites that experienced a minimum of 100 crashes per year. Given the fact that the minimum sample size is a function of sample variance, this recommendation is clearly questionable as roadway characteristics of different roadway types are likely to have different levels of homogeneity. This research used Florida data to determine the minimum sample sizes to estimate reliable calibration factors for the following three facility types: rural two-lane roads, rural multilane highways, and urban and suburban arterials. The analysis was based on the data collected from more than 7,000 miles of segments and more than 1,000 intersections in Florida. The minimum sample size was determined such that there is a high probability that the calibration factor estimated from a sample is within 5% to 10% of the actual calibration factor calculated from the entire data set. For all the site subtypes, it was found that the minimum sample size of 30 to 50 sites, as recommended by the HSM, is insufficient to achieve the desired accuracy. Moreover, the sample sizes required for estimated calibration factors to lie within 5% of the actual calibration factors is almost double the sample sizes required for estimated calibration factors to lie within 10% of the actual calibration factors. The results also showed that the generalized one-size-fits-all approach of using a sample size of 30 to 50 sites is not appropriate as different facility types require different sample sizes depending on several factors, such as the extent of data variability, population size, crash experience, and so on, to estimate reliable calibration factors.

  • Research Article
  • Cite Count Icon 46
  • 10.1093/jee/99.2.568
Sampling Methods, Dispersion Patterns, and Fixed Precision Sequential Sampling Plans for Western Flower Thrips (Thysanoptera: Thripidae) and Cotton Fleahoppers (Hemiptera: Miridae) in Cotton
  • Apr 1, 2006
  • Journal of Economic Entomology
  • M N Parajulee + 2 more

A 2-yr field study was conducted to examine the effectiveness of two sampling methods (visual and plant washing techniques) for western flower thrips, Frankliniella occidentalis (Pergande), and five sampling methods (visual, beat bucket, drop cloth, sweep net, and vacuum) for cotton fleahopper, Pseudatomoscelis seriatus (Reuter), in Texas cotton, Gossypium hirsutum (L.), and to develop sequential sampling plans for each pest. The plant washing technique gave similar results to the visual method in detecting adult thrips, but the washing technique detected significantly higher number of thrips larvae compared with the visual sampling. Visual sampling detected the highest number of fleahoppers followed by beat bucket, drop cloth, vacuum, and sweep net sampling, with no significant difference in catch efficiency between vacuum and sweep net methods. However, based on fixed precision cost reliability, the sweep net sampling was the most cost-effective method followed by vacuum, beat bucket, drop cloth, and visual sampling. Taylor’s Power Law analysis revealed that the field dispersion patterns of both thrips and fleahoppers were aggregated throughout the crop growing season. For thrips management decision based on visual sampling (0.25 precision), 15 plants were estimated to be the minimum sample size when the estimated population density was one thrips per plant, whereas the minimum sample size was nine plants when thrips density approached 10 thrips per plant. The minimum visual sample size for cotton fleahoppers was 16 plants when the density was one fleahopper per plant, but the sample size decreased rapidly with an increase in fleahopper density, requiring only four plants to be sampled when the density was 10 fleahoppers per plant. Sequential sampling plans were developed and validated with independent data for both thrips and cotton fleahoppers.

  • Research Article
  • Cite Count Icon 24
  • 10.1603/0022-0493-99.2.568
Sampling Methods, Dispersion Patterns, and Fixed Precision Sequential Sampling Plans for Western Flower Thrips (Thysanoptera: Thripidae) and Cotton Fleahoppers (Hemiptera: Miridae) in Cotton
  • Apr 1, 2006
  • Journal of Economic Entomology
  • M N Parajulee + 2 more

A 2-yr field study was conducted to examine the effectiveness of two sampling methods (visual and plant washing techniques) for western flower thrips, Frankliniella occidentalis (Pergande), and five sampling methods (visual, beat bucket, drop cloth, sweep net, and vacuum) for cotton fleahopper, Pseudatomoscelis seriatus (Reuter), in Texas cotton, Gossypium hirsutum (L.), and to develop sequential sampling plans for each pest. The plant washing technique gave similar results to the visual method in detecting adult thrips, but the washing technique detected significantly higher number of thrips larvae compared with the visual sampling. Visual sampling detected the highest number of fleahoppers followed by beat bucket, drop cloth, vacuum, and sweep net sampling, with no significant difference in catch efficiency between vacuum and sweep net methods. However, based on fixed precision cost reliability, the sweep net sampling was the most cost-effective method followed by vacuum, beat bucket, drop cloth, and visual sampling. Taylor's Power Law analysis revealed that the field dispersion patterns of both thrips and fleahoppers were aggregated throughout the crop growing season. For thrips management decision based on visual sampling (0.25 precision), 15 plants were estimated to be the minimum sample size when the estimated population density was one thrips per plant, whereas the minimum sample size was nine plants when thrips density approached 10 thrips per plant. The minimum visual sample size for cotton fleahoppers was 16 plants when the density was one fleahopper per plant, but the sample size decreased rapidly with an increase in fleahopper density, requiring only four plants to be sampled when the density was 10 fleahoppers per plant. Sequential sampling plans were developed and validated with independent data for both thrips and cotton fleahoppers.

  • Research Article
  • Cite Count Icon 4
  • 10.1002/pc.27580
Analysis of sample size for friction test of short fiber reinforced polymer composites
  • Jul 25, 2023
  • Polymer Composites
  • Bo‐Wen Guan + 3 more

According to the recommendations of literature and standards, the minimum sample size for the friction test of thermoplastics is usually 3–5. While this size is valid for the friction test of neat polymer resins with homogeneous structures, its applicability to short fiber reinforced polymer composites with inhomogeneous structures remains unclear because the uneven distribution of discontinuous fibers or fillers in the matrix leads to the performance of the composites being dispersed. In this work, as a typical example, 100 friction samples made of short carbon fiber reinforced polyetherimide (SCF/PEI) composite with a typical content of 10 vol% SCFs are prepared, and their friction coefficients are measured. The maximum likelihood estimation and goodness of fit methods are employed to assess the fitting degree between the overall friction coefficient results and five mathematical distributions. The results show that the overall friction coefficient data distribution is in good agreement with the Lognormal, Generalized extreme value (GEV), and Burr distributions, respectively. Based on the fitting result, the minimum sample size is analyzed by Monte Carlo Simulation. It shows that the minimum sample size strongly depends on the property dispersion parameter (COV: coefficient of variation). Based on this parameter, the minimum sample size is given. In a practical test, the performance dispersion degree of a composite cannot be known in advance, hence a sample size determination method is proposed to yield a reasonable performance description parameter by applying the results of this work, and this method is validated by the overall test data.Highlights Proposed a sample size determination method for practical friction tests. The minimum sample size for the friction test is yielded statistically. The minimum sample size strongly depends on a property dispersion parameter.

  • Research Article
  • Cite Count Icon 177
  • 10.2427/12267
Guidelines of the minimum sample size requirements for Kappa agreement test
  • Mar 29, 2022
  • Epidemiology, Biostatistics, and Public Health
  • Mohamad Adam Bujang + 1 more

Background and aims: To estimate sample size for Cohen’s kappa agreement test can be challenging especially when it is expected that the true marginal rating frequencies are not the same. This study aims to present the tables that display a minimum sample size determination for an agreement test when certain assumptions are hold. Method: We adopted the sample size formula provided by Flack and colleagues (1988) to calculate the minimum sample sizes required using PASS software. The power is pre-specified to be at least 80% and the alpha to be less than 0.05. The effect sizes were derived from several pre-specified estimates such as the pattern of the true marginal rating frequencies and the difference between the two kappa coefficients in the hypothesis testing. Results: When the true marginal rating frequencies are the same, the minimum sample size determination can range from 2 to 698 depending on the actual value of the effect size. When the true marginal rating frequencies are not the same, then the majority of the minimum sample size required for this condition is more than double than that required sample size when the true marginal rating frequencies are the same. Conclusion: Concerning that the sample size formula could produce a very extreme small sample size, therefore, the determination of K1 and K2 should be based on reasonable estimates. We recommend for all sample size determinations for Cohen’s kappa agreement test, the true marginal rating frequencies can be assumed the same. Otherwise, it will be necessary to multiple the estimated minimum sample size by two to accommodate if the true marginal rating frequencies is not the same.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 12
  • 10.1186/s13595-023-01209-4
Effect of sample size on the estimation of forest inventory attributes using airborne LiDAR data in large-scale subtropical areas
  • Oct 20, 2023
  • Annals of Forest Science
  • Chungan Li + 4 more

Key messageSample size (number of plots) may significantly affect the accuracy of forest attribute estimations using airborne LiDAR data in large-scale subtropical areas. In general, the accuracy of all models improves with increasing sample size. However, the improvement in estimation accuracy varies across forest attributes and forest types. Overall, a larger sample size is required to estimate the stand volume (VOL), while a smaller sample size is required to estimate the mean diameter at breast height (DBH). Broad-leaved forests require a smaller sample size than Chinese fir forests.ContextSample size is an essential factor affecting the cost of LiDAR-assisted forest resource inventory. Therefore, investigating the minimum sample size required to achieve acceptable accuracy for airborne LiDAR-based forest attribute estimation can help improve cost efficiency and optimize technical schemes.AimsThe aims were to assess the optimal sample size to estimate the VOL, basal area, mean height, and DBH in stands dominated by Cunninghamia lanceolate, Pinus massoniana, Eucalyptus spp., and other broad-leaved species in a large subtropical area using airborne LiDAR data.MethodsStatistical analyses were performed on the differences in LiDAR metrics between different sample sizes and the total number of plots, as well as on the field-measured attributes. The relative root mean square error (rRMSE) and the determination coefficient (R2) of multiplicative power models with different sample sizes were compared. The logistic regression between the coefficient of variation of the rRMSE and the sample size was established, and the minimum sample size was determined using a threshold of less than 10% for the coefficient of variation.ResultsAs the sample sizes increased, we found a decrease in the mean rRMSE and an increase in the mean R2, as well as a decrease in the standard deviation of the LiDAR metrics and field-measured attributes. Sample sizes for Chinese fir, pine, eucalyptus, and broad-leaved forests should be over 110, 80, 85, and 60, respectively, in a practical airborne LiDAR-based forest inventory.ConclusionThe accuracy of all forest attribute estimations improved as the sample size increased across all forest types, which could be attributed to the decreasing variations of both LiDAR metrics and field-measured attributes.

  • Research Article
  • Cite Count Icon 23
  • 10.1016/j.hrmr.2016.09.012
Robustness of statistical inferences using linear models with meta-analytic correlation matrices
  • Oct 5, 2016
  • Human Resource Management Review
  • Patrick J Rosopa + 1 more

Robustness of statistical inferences using linear models with meta-analytic correlation matrices

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant