Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A solution to minimum sample size for regressions

  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Regressions and meta-regressions are widely used to estimate patterns and effect sizes in various disciplines. However, many biological and medical analyses use relatively low sample size (N), contributing to concerns on reproducibility. What is the minimum N to identify the most plausible data pattern using regressions? Statistical power analysis is often used to answer that question, but it has its own problems and logically should follow model selection to first identify the most plausible model. Here we make null, simple linear and quadratic data with different variances and effect sizes. We then sample and use information theoretic model selection to evaluate minimum N for regression models. We also evaluate the use of coefficient of determination (R2) for this purpose; it is widely used but not recommended. With very low variance, both false positives and false negatives occurred at N < 8, but data shape was always clearly identified at N ≥ 8. With high variance, accurate inference was stable at N ≥ 25. Those outcomes were consistent at different effect sizes. Akaike Information Criterion weights (AICc wi) were essential to clearly identify patterns (e.g., simple linear vs. null); R2 or adjusted R2 values were not useful. We conclude that a minimum N = 8 is informative given very little variance, but minimum N ≥ 25 is required for more variance. Alternative models are better compared using information theory indices such as AIC but not R2 or adjusted R2. Insufficient N and R2-based model selection apparently contribute to confusion and low reproducibility in various disciplines. To avoid those problems, we recommend that research based on regressions or meta-regressions use N ≥ 25.

Similar Papers
  • Research Article
  • Cite Count Icon 84
  • 10.1016/j.fishres.2009.03.016
Gear selectivity and sample size effects on growth curve selection in shark age and growth studies
  • Apr 9, 2009
  • Fisheries Research
  • James T Thorson + 1 more

Gear selectivity and sample size effects on growth curve selection in shark age and growth studies

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 8
  • 10.1007/s00180-016-0690-2
On the impact of model selection on predictor identification and parameter inference
  • Oct 22, 2016
  • Computational Statistics
  • Ruth M Pfeiffer + 2 more

We assessed the ability of several penalized regression methods for linear and logistic models to identify outcome-associated predictors and the impact of predictor selection on parameter inference for practical sample sizes. We studied effect estimates obtained directly from penalized methods (Algorithm 1), or by refitting selected predictors with standard regression (Algorithm 2). For linear models, penalized linear regression, elastic net, smoothly clipped absolute deviation (SCAD), least angle regression and LASSO had a low false negative (FN) predictor selection rates but false positive (FP) rates above 20 % for all sample and effect sizes. Partial least squares regression had few FPs but many FNs. Only relaxo had low FP and FN rates. For logistic models, LASSO and penalized logistic regression had many FPs and few FNs for all sample and effect sizes. SCAD and adaptive logistic regression had low or moderate FP rates but many FNs. 95 % confidence interval coverage of predictors with null effects was approximately 100 % for Algorithm 1 for all methods, and 95 % for Algorithm 2 for large sample and effect sizes. Coverage was low only for penalized partial least squares (linear regression). For outcome-associated predictors, coverage was close to 95 % for Algorithm 2 for large sample and effect sizes for all methods except penalized partial least squares and penalized logistic regression. Coverage was sub-nominal for Algorithm 1. In conclusion, many methods performed comparably, and while Algorithm 2 is preferred to Algorithm 1 for estimation, it yields valid inference only for large effect and sample sizes.

  • Research Article
  • Cite Count Icon 1
  • 10.1525/elementa.2025.00020
Assessing the performance of emerging and existing continuous monitoring solutions under a single-blind controlled testing protocol
  • Sep 29, 2025
  • Elem Sci Anth
  • Fancy Cheptonui + 7 more

Continuous monitoring (CM) solutions can facilitate faster detection and repair of emissions compared to traditional survey methods. This study tested 13 CM solutions over 12 weeks using single-blind controlled testing. Controlled release rates ranged from 0.08 to 6.75 kg CH4 h−1 and lasted 18 min to 8 h. Six solutions demonstrated 90% method detection limits (DL90s) ranging from 0.5 [0.3, 0.6] kg CH4 h−1 to 6.7 [5.9, 8.0] kg CH4 h−1. Of the 6 solutions, 5 had low False Positive (FP) rates (7.8%–18.9%), and 4 had low False Negative (FN) rates (8.0%–34.1%). Similar to Ilonze et al., the results show that the tested solutions balance method sensitivity with low FP and FN rates. All scanning/imaging solutions achieved high localization precision and accuracy (≥40%) at the equipment unit level. Single quantification estimates exhibited high relative quantification errors, ranging from 33 [0.9, 66]%, 95% confidence interval (CI) to 1326 [1003, 1648]%, 95% CI for small emissions (between 0.1 and 1 kg CH4 h−1) and from 3 [−20, 26]%, 95% CI to 3578 [−2832, 9988]%, 95% CI for large emissions (&amp;gt;1 kg CH4 h−1). The mean detection time for all solutions ranged from 5 h to 5 days. Relative to previous studies, errors in quantification estimates decreased, as did FN and FP rates, with improved DL90s for 2 of the 4 retested solutions. However, the mean detection times increased for 2 solutions, remained constant for one solution, and decreased for 1 of the 4 retested solutions. These findings highlight that continuous, rigorous testing enhances solution performance, with notable improvements observed across multiple testing programs using the same test protocol.

  • Research Article
  • Cite Count Icon 14
  • 10.1142/9789813207813_0049
IDENTIFYING GENETIC ASSOCIATIONS WITH VARIABILITY IN METABOLIC HEALTH AND BLOOD COUNT LABORATORY VALUES: DIVING INTO THE QUANTITATIVE TRAITS BY LEVERAGING LONGITUDINAL DATA FROM AN EHR.
  • Nov 22, 2016
  • Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
  • Shefali S Verma + 14 more

A wide range of patient health data is recorded in Electronic Health Records (EHR). This data includes diagnosis, surgical procedures, clinical laboratory measurements, and medication information. Together this information reflects the patient's medical history. Many studies have efficiently used this data from the EHR to find associations that are clinically relevant, either by utilizing International Classification of Diseases, version 9 (ICD-9) codes or laboratory measurements, or by designing phenotype algorithms to extract case and control status with accuracy from the EHR. Here we developed a strategy to utilize longitudinal quantitative trait data from the EHR at Geisinger Health System focusing on outpatient metabolic and complete blood panel data as a starting point. Comprehensive Metabolic Panel (CMP) as well as Complete Blood Counts (CBC) are parts of routine care and provide a comprehensive picture from high level screening of patients' overall health and disease. We randomly split our data into two datasets to allow for discovery and replication. We first conducted a genome-wide association study (GWAS) with median values of 25 different clinical laboratory measurements to identify variants from Human Omni Express Exome beadchip data that are associated with these measurements. We identified 687 variants that associated and replicated with the tested clinical measurements at p<5×10-08. Since longitudinal data from the EHR provides a record of a patient's medical history, we utilized this information to further investigate the ICD-9 codes that might be associated with differences in variability of the measurements in the longitudinal dataset. We identified low and high variance patients by looking at changes within their individual longitudinal EHR laboratory results for each of the 25 clinical lab values (thus creating 50 groups - a high variance and a low variance for each lab variable). We then performed a PheWAS analysis with ICD-9 diagnosis codes, separately in the high variance group and the low variance group for each lab variable. We found 717 PheWAS associations that replicated at a p-value less than 0.001. Next, we evaluated the results of this study by comparing the association results between the high and low variance groups. For example, we found 39 SNPs (in multiple genes) associated with ICD-9 250.01 (Type-I diabetes) in patients with high variance of plasma glucose levels, but not in patients with low variance in plasma glucose levels. Another example is the association of 4 SNPs in UMOD with chronic kidney disease in patients with high variance for aspartate aminotransferase (discovery p-value: 8.71×10-09 and replication p-value: 2.03×10-06). In general, we see a pattern of many more statistically significant associations from patients with high variance in the quantitative lab variables, in comparison with the low variance group across all of the 25 laboratory measurements. This study is one of the first of its kind to utilize quantitative trait variance from longitudinal laboratory data to find associations among genetic variants and clinical phenotypes obtained from an EHR, integrating laboratory values and diagnosis codes to understand the genetic complexities of common diseases.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 42
  • 10.1007/s10940-020-09467-5
Police Legitimacy and the Norm to Cooperate: Using a Mixed Effects Location-Scale Model to Estimate the Strength of Social Norms at a Small Spatial Scale
  • Jul 21, 2020
  • Journal of Quantitative Criminology
  • Jonathan Jackson + 5 more

ObjectivesTest whether cooperation with the police can be modelled as a place-based norm that varies in strength from one neighborhood to the next. Estimate whether perceived police legitimacy predicts an individual’s willingness to cooperate in weak-norm neighborhoods, but not in strong-norm neighborhoods where most people are either willing or unwilling to cooperate, irrespective of their perceptions of police legitimacy.MethodsA survey of 1057 individuals in 98 relatively high-crime English neighborhoods defined at a small spatial scale measured (a) willingness to cooperate using a hypothetical crime vignette and (b) legitimacy using indicators of normative alignment between police and citizen values. A mixed-effects, location-scale model estimated the cluster-level mean and cluster-level variance of willingness to cooperate as a neighborhood-level latent variable. A cross-level interaction tested whether legitimacy predicts individual-level willingness to cooperate only in neighborhoods where the norm is weak.ResultsWillingness to cooperate clustered strongly by neighborhood. There were neighborhoods with (1) high mean and low variance, (2) high mean and high variance, (3) (relatively) low mean and low variance, and (4) (relatively) low mean and high variance. Legitimacy was only a positive predictor of cooperation in neighborhoods that had a (relatively) low mean and high variance. There was little variance left to explain in neighborhoods where the norm was strong.ConclusionsFindings support a boundary condition of procedural justice theory: namely, that cooperation can be modelled as a place-based norm that varies in strength from neighborhood to neighborhood and that legitimacy only predicts an individual’s willingness to cooperate in neighborhoods where the norm is relatively weak.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 31
  • 10.1371/journal.pone.0146491
Model Selection in Historical Research Using Approximate Bayesian Computation.
  • Jan 5, 2016
  • PLOS ONE
  • Xavier Rubio-Campillo

Formal Models and HistoryComputational models are increasingly being used to study historical dynamics. This new trend, which could be named Model-Based History, makes use of recently published datasets and innovative quantitative methods to improve our understanding of past societies based on their written sources. The extensive use of formal models allows historians to re-evaluate hypotheses formulated decades ago and still subject to debate due to the lack of an adequate quantitative framework. The initiative has the potential to transform the discipline if it solves the challenges posed by the study of historical dynamics. These difficulties are based on the complexities of modelling social interaction, and the methodological issues raised by the evaluation of formal models against data with low sample size, high variance and strong fragmentation.Case StudyThis work examines an alternate approach to this evaluation based on a Bayesian-inspired model selection method. The validity of the classical Lanchester’s laws of combat is examined against a dataset comprising over a thousand battles spanning 300 years. Four variations of the basic equations are discussed, including the three most common formulations (linear, squared, and logarithmic) and a new variant introducing fatigue. Approximate Bayesian Computation is then used to infer both parameter values and model selection via Bayes Factors.ImpactResults indicate decisive evidence favouring the new fatigue model. The interpretation of both parameter estimations and model selection provides new insights into the factors guiding the evolution of warfare. At a methodological level, the case study shows how model selection methods can be used to guide historical research through the comparison between existing hypotheses and empirical evidence.

  • Research Article
  • Cite Count Icon 47
  • 10.1037/met0000145
A comparison of Bayesian and frequentist model selection methods for factor analysis models.
  • Jun 1, 2017
  • Psychological Methods
  • Zhao-Hua Lu + 2 more

We compare the performances of well-known frequentist model fit indices (MFIs) and several Bayesian model selection criteria (MCC) as tools for cross-loading selection in factor analysis under low to moderate sample sizes, cross-loading sizes, and possible violations of distributional assumptions. The Bayesian criteria considered include the Bayes factor (BF), Bayesian Information Criterion (BIC), Deviance Information Criterion (DIC), a Bayesian leave-one-out with Pareto smoothed importance sampling (LOO-PSIS), and a Bayesian variable selection method using the spike-and-slab prior (SSP; Lu, Chow, & Loken, 2016). Simulation results indicate that of the Bayesian measures considered, the BF and the BIC showed the best balance between true positive rates and false positive rates, followed closely by the SSP. The LOO-PSIS and the DIC showed the highest true positive rates among all the measures considered, but with elevated false positive rates. In comparison, likelihood ratio tests (LRTs) are still the preferred frequentist model comparison tool, except for their higher false positive detection rates compared to the BF, BIC and SSP under violations of distributional assumptions. The root mean squared error of approximation (RMSEA) and the Tucker-Lewis index (TLI) at the conventional cut-off of approximate fit impose much more stringent "penalties" on model complexity under conditions with low cross-loading size, low sample size, and high model complexity compared with the LRTs and all other Bayesian MCC. Nevertheless, they provided a reasonable alternative to the LRTs in cases where the models cannot be readily constructed as nested within each other. (PsycINFO Database Record

  • Research Article
  • Cite Count Icon 36
  • 10.1016/j.ympev.2010.08.022
The evolution of autodigestion in the mushroom family Psathyrellaceae (Agaricales) inferred from Maximum Likelihood and Bayesian methods
  • Aug 27, 2010
  • Molecular Phylogenetics and Evolution
  • László G Nagy + 5 more

The evolution of autodigestion in the mushroom family Psathyrellaceae (Agaricales) inferred from Maximum Likelihood and Bayesian methods

  • Research Article
  • Cite Count Icon 5
  • 10.3390/sym14102159
Crowd Density Estimation in Spatial and Temporal Distortion Environment Using Parallel Multi-Size Receptive Fields and Stack Ensemble Meta-Learning
  • Oct 15, 2022
  • Symmetry
  • Addis Abebe Assefa + 3 more

The estimation of crowd density is crucial for applications such as autonomous driving, visual surveillance, crowd control, public space planning, and warning visually distracted drivers prior to an accident. Having strong translational, reflective, and scale symmetry, models for estimating the density of a crowd yield an encouraging result. However, dynamic scenes with perspective distortions and rapidly changing spatial and temporal domains still present obstacles. The main reasons for this are the dynamic nature of a scene and the difficulty of representing and incorporating the feature space of objects of varying sizes into a prediction model. To overcome the aforementioned issues, this paper proposes a parallel multi-size receptive field units framework that leverages the majority of the CNN layer’s features, allowing for the representation and participation in the model prediction of the features of objects of all sizes. The proposed method utilizes features generated from lower to higher layers. As a result, different object scales can be handled at different framework depths, and various environmental densities can be estimated. However, the inclusion of the vast majority of layer features in the prediction model has a number of negative effects on the prediction’s outcome. Asymmetric non-local attention and the channel weighting module of a feature map are proposed to handle noise and background details and re-weight each channel to make it more sensitive to important features while ignoring irrelevant ones, respectively. While the output predictions of some layers have high bias and low variance, those of other layers have low bias and high variance. Using stack ensemble meta-learning, we combine individual predictions made with lower-layer features and higher-layer features to improve prediction while balancing the tradeoff between bias and variance. The UCF CC 50 dataset and the ShanghaiTech dataset have both been subjected to extensive testing. The results of the experiments indicate that the proposed method is effective for dense distributions and objects of various sizes.

  • Research Article
  • Cite Count Icon 1
  • 10.1214/25-ejs2353
Road traffic estimation and distribution-based route selection
  • Jan 1, 2025
  • Electronic Journal of Statistics
  • Rens Kamphuis + 2 more

In route selection problems, the driver’s personal preferences will determine whether she prefers a route with a travel time that has a relatively low mean and high variance over one that has relatively high mean and low variance. In practice, however, such risk aversion issues are often ignored, in that a route is selected based on a single-criterion Dijkstra-type algorithm. In addition, the routing decision typically does not take into account the uncertainty in the estimates of the travel time’s distribution. This paper aims at resolving both issues by setting up a framework for travel time estimation. In our framework, the underlying road network is represented as a graph. Each edge is subdivided into multiple smaller pieces, so as to naturally model the statistical similarity between road pieces that are spatially nearby. Relying on a Bayesian approach, we construct an estimator for the joint per-edge travel time distribution, thus also providing us with an uncertainty quantification of our estimates. Our machinery relies on establishing limit theorems, making the resulting estimation procedure robust in the sense that it does not hinge on any distributional properties but instead on a working model. We present an extensive set of numerical experiments that demonstrate the validity of the estimation procedure and the use of the distributional estimates in the context of data-driven route selection.

  • Conference Article
  • Cite Count Icon 1
  • 10.1109/nafips.2004.1337396
Efficient parameter selection for system identification
  • Jan 1, 2004
  • IEEE Annual Meeting of the Fuzzy Information, 2004. Processing NAFIPS '04.
  • A Ghodsi

There are a wide variety of techniques for system identification. A common critical issue in all of these techniques is selecting the appropriate complexity. In particular, system identification algorithms based on two popular clustering techniques i.e., subtractive clustering and c-mean (FCM) clustering, require that the number of underlying partitions be selected ahead of time. That is, if they do not automatically choose the appropriate complexity to model the data. They only find the best-fit model for a given complexity. A model with an overly restricted complexity gives poor predictions on new data, since the model has too little flexibility (yielding high bias and low variance). By contrast, a model with too complexity also gives poor generalization performance since it is too flexible and fits too much of the noise on the training data (yielding low bias but high variance). Bias and variance are complementary quantities, and it is necessary to assign the complexity optimally in order to achieve the best compromise between them. In this paper we propose a general criterion for choosing an appropriate complexity based on a simple resampling approach. We achieve this by deriving a generalization of analytic form of leave-one-out cross-validation risk estimator, which we can then use to determine the optimal complexity of a model in a given data set.

  • Research Article
  • Cite Count Icon 57
  • 10.1016/s0893-6080(03)00118-7
Automatic basis selection techniques for RBF networks
  • May 15, 2003
  • Neural Networks
  • Ali Ghodsi + 1 more

Automatic basis selection techniques for RBF networks

  • Research Article
  • Cite Count Icon 30
  • 10.1007/s00330-015-3812-2
Radiologic-pathologic analysis of quantitative 3D tumour enhancement on contrast-enhanced MR imaging: a study of ROI placement.
  • May 21, 2015
  • European Radiology
  • Arun Chockalingam + 9 more

To investigate the influence of region-of-interest (ROI) placement on 3D tumour enhancement [Quantitative European Association for the Study of the Liver (qEASL)] in hepatocellular carcinoma (HCC) patients treated with transcatheter arterial chemoembolization (TACE). Phase 1: 40 HCC patients had nine ROIs placed by one reader using systematic techniques (3 ipsilateral to the lesion, 3 contralateral to the lesion, and 3 dispersed throughout the liver) and qEASL variance was measured. Intra-class correlations were computed. Phase 2: 15 HCC patients with histosegmentation were selected. Six ROIs were systematically placed by AC (3 ROIs ipsilateral and 3 ROIs contralateral to the lesion). Three ROIs were placed by 2 radiologists. qEASL values were compared to histopathology by Pearson's correlation, linear regression, and median difference. Phase 1: The dispersed method (abandoned in phase 2) had low consistency and high variance. Phase 2: qEASL correlated strongly with pathology in systematic methods [Pearson's correlation coefficient = 0.886 (ipsilateral) and 0.727 (contralateral)] and in clinical methods (0.625 and 0.879). However, ipsilateral placement matched best with pathology (median difference: 5.4 %; correlation: 0.89; regression CI: [0.904, 0.1409]). qEASL is a robust method with comparable values among tested placements. Ipsilateral placement showed high consistency and better pathological correlation. Ipsilateral and contralateral ROI placement produces high consistency and low variance. Both ROI placement methods produce qEASL values that correlate well with histopathology. Ipsilateral ROI placement produces best correlation to pathology along with high consistency.

  • Research Article
  • Cite Count Icon 23
  • 10.1897/09-032.1
Challenges in current adult fish laboratory reproductive tests: Suggestions for refinement using a mummichog (Fundulus heteroclitus) case study
  • Nov 1, 2009
  • Environmental Toxicology and Chemistry
  • Thus Bosker + 2 more

Concerns about screening endocrine-active contaminants have led to the development of a number of short-term fish reproductive tests. A review conducted of 62 published adult fish reproductive papers using various fish species found low samples sizes (mean of 5.7 replicates with a median of 5 replicates) and high variance (an average coefficient of variance of 43.8%). The high variances and low sample sizes allow only relatively large differences to be detected with the current protocols; the average significant difference detected was a 68.7% reduction in egg production, while only differences above 50% were detected with confidence. This result indicates low power to detect more subtle differences and a high probability of type II errors in interpretation. The present study identifies several ways to increase the power of the adult fish reproductive test in the mummichog (Fundulus heteroclitus). By identifying the peak timing of egg production (before and after the new moon), extending the duration of the experiment (increased from 7 to 14 d), and determining that a sample size of eight replicate tanks per treatment accurately predicts variance in the sample population (based on pre-exposure variation calculations of replicate tanks), the power of the test has been significantly increased. The present study demonstrates that weaknesses in the current adult fish reproductive tests can easily be addressed by focusing on improved understanding of the reproductive behavior of the test species and developing study designs that include calculating desired variability levels and increasing replicates.

  • Research Article
  • Cite Count Icon 10
  • 10.2134/jpa1994.0211
Increasing Plot Length Reduces Experimental Error of On-Farm Tests
  • Apr 1, 1994
  • Journal of Production Agriculture
  • Stewart B Wuest + 6 more

Farmer conducted on-farm research is an effective tool for development of crop management practices. The randomized complete-block experimental design is being used in on-farm tests, with blocks consisting of two or more long, narrow, side-by-side plots. This study examined the relationship between plot length and experimental error, and assesses the probable statistical outcome of on-farm tests performed in the dryland region of the Pacific Northwest, USA. Fourteen trials were conducted in wheat (Triticum aestivum L.) and barley (Hordeum vulgare L.) fields to measure yield variance of combine-width plots ranging in length from 250 to 1500 ft. The relationship between plot length and variance for each site followed a logarithmic decay model (average r2 = 0.88). Variance declined rapidly as plot length increased from 250 to 750 ft at most sites. Averaging the ten least variable sites, the LSD (0.05) with three degrees of freedom declined from 6.5 to 2.6 bu/acre as plot length increased from 250 to 1500 ft. At the same 10 sites, power for mean separation (α = 0.05) of treatments with 4 bu/acre true difference was >0.80 with six replications and 750 ft plot length, or four replications and 1250 ft plot length. With adequate replication and plot length, on-farm tests can be designed for highly variable dryland regions with good control of experimental error. Research Question The use of side-by-side, combine-width experimental plots in farmer-conducted on-farm tests is increasing. Information on how plot length affects the success of these tests is lacking. It is also desirable to know what level of precision to expect from on-farm tests performed in highly variable dryland cereal production areas. This study investigated the relationship between the length of combine-wide, side-by-side plots and experimental error under the dryland grain production conditions of the Pacific Northwest. Literature Summary Researchers studying the performance of on-farm tests have found that the randomized complete block design used in the Midwest produces results comparable in precision to research station small plot experiments. These designs use 1200 to 1300 ft long plots 20 to 40 ft wide, and primarily involve corn or soybean. The performance of on-farm tests under dryland small grain production conditions has not been evaluated. There has also been little research to date that can be used to recommend minimum, maximum, or optimum plot lengths. Study Description Fourteen uniformity trials were harvested in commercial wheat and barley fields in Washington, Idaho, and Oregon. A uniformity trial measures natural variability between plots by harvesting plots where no treatments have been placed. Side-by-side strips 1500 ft long were harvested in 250 ft segments to allow recombination of the data into plots of different lengths. The grain yield data (bu/acre) were analyzed to determine variance between pairs of side-by-side plots. Ten of fourteen sites had a variance of <5 at plot lengths of 1500 ft, and were classified as low variance. The remaining four sites ranged from 6 to 22 (high variance). Figure 1 shows least significant differences at α = 0.05 (LSD 0.05) for an individual on-farm test with two treatments and four replications based on average variances for the low and high variance groups. Applied Questions How long should on-farm test plots be? Figure 1Open in figure viewerPowerPoint The effect of increasing plot length on LSD values at α = 0.05 is shown for experiments with different numbers of replications (reps). (A) is based upon data from the 10 fields with low variance and (B) from the remaining four fields with high variance. Note that the scales on the vertical axes differ. In most fields there was a large decrease in variance as plot length increased. Therefore, on-farm tests will produce more reliable results as plot length increases from 250 to 750 ft or more. In some very uniform fields even short plots will have acceptably low variability, but in every field measured, variability decreased as plot length increased to 1500 ft. To ensure the best results, we recommend that plots be as long as is practical. How does replication affect precision? The variability encountered in field experiments makes replication the key to a successful test. Figure 1 shows how LSD (0.05) decreases when replications are added. A low LSD is important because it allows detection of smaller differences between the performance of the treatments, or if there is no difference, it allows a high confidence that the treatments do not perform differently. On-farm tests can provide valuable information to farmers and researchers. Small differences can be detected with a high degree of confidence in most fields with four or more replications of 1000 ft or longer side-by-side plots.

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface