Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

The arcsine is asinine: the analysis of proportions in ecology

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

The arcsine square root transformation has long been standard procedure when analyzing proportional data in ecology, with applications in data sets containing binomial and non-binomial response variables. Here, we argue that the arcsine transform should not be used in either circumstance. For binomial data, logistic regression has greater interpretability and higher power than analyses of transformed data. However, it is important to check the data for additional unexplained variation, i.e., overdispersion, and to account for it via the inclusion of random effects in the model if found. For non-binomial data, the arcsine transform is undesirable on the grounds of interpretability, and because it can produce nonsensical predictions. The logit transformation is proposed as an alternative approach to address these issues. Examples are presented in both cases to illustrate these advantages, comparing various methods of analyzing proportions including untransformed, arcsine- and logit-transformed linear models and logistic regression (with or without random effects). Simulations demonstrate that logistic regression usually provides a gain in power over other methods.

Similar Papers
  • Research Article
  • Cite Count Icon 10
  • 10.1653/024.096.0361
Logistic Regression is a better Method of Analysis Than Linear Regression of Arcsine Square Root Transformed Proportional Diapause Data ofPieris melete(Lepidoptera: Pieridae)
  • Sep 1, 2013
  • Florida Entomologist
  • P J Shi + 2 more

Temperature and day-length are considered to be the 2 important factors that can significantly affect insect diapause, which is a typical proportional dataset. In the previous studies, the method of arcsine square root transformation is widely used to analyze the effect of temperature or day-length or their joint effects on diapause in insects. However, this method has many limitations, for example, the proportional data should be normally distributed. The logistic regression in generalized additive models is a promising method for analyzing the effects of temperature and day-length on diapause. Compared to the arcsine square root transformation method, this method does not require normal distribution of proportional diapause data. The logistic regression also provides better goodness-of-fit by using the non-parametric fitting technique. In this report, we used the diapause data of Pieris melete (Xiao et al. 2012) to compare the fitted results of the logistic regression in generalized additive models with arcsine square root transformation. We found that the logistic regression in generalized additive models is better than linear regression of arcsine square root transformed data in following ways: (1) reasonable predictions about diapause ranging from 0 to 1 can be made without transforming the proportional data; (2) non-linear effects of temperature and day-length on diapause can be determined; (3) the goodness-of-fit can be substantially improved. View this article in BioOne

  • Research Article
  • Cite Count Icon 16
  • 10.1002/bdr2.1755
Analysis of proportional data in reproductive and developmental toxicity studies: Comparison of sensitivities of logit transformation, arcsine square root transformation, and nonparametric analysis.
  • Jul 31, 2020
  • Birth Defects Research
  • Paul I Feder + 4 more

In developmental and reproductive toxicity studies, analysis of litter-based binary endpoints (e.g., incidence of malformed fetuses) is complex in that littermates often are not entirely independent of one another. It is well established that the litter, not the individual fetus, is the proper independent experimental unit in statistical analysis. Accordingly, analysis is often based on the proportion affected per litter and the litter proportions are analyzed as continuous data. Because these proportional data generally do not meet assumptions of symmetry or normality, data are typically analyzed by nonparametric methods, arcsine square root transformation, or logit transformation. We conducted power calculations to compare different approaches (nonparametric, arcsine square root-transformed, logit-transformed, untransformed) for analyzing litter-based proportional data. A reproductive toxicity study with a control and one treated group provided data for two endpoints: prenatal loss, and fertility by in utero insemination (IUI). Type 1 error and power were estimated by 10,000 simulations based on two-sample one-tailed t tests with varying numbers of litters per group. To further compare the different approaches, we conducted additional analyses with shifted mean proportions to produce illustrative scenarios. Analyses based on logit-transformed proportions had greater power than those based on untransformed or arcsine square root-transformed proportions, or nonparametric procedures. The logit transformation is preferred to the other approaches considered when making inferences concerning litter-based proportional endpoints, particularly with skewed distributions. The improved performance of the logit transformation becomes increasingly pronounced as the response proportions are increasingly close to the boundaries of the parameter space.

  • Front Matter
  • 10.1351/goldbook.14445
Arcsine square root transformation
  • Oct 26, 2025

Citation: 'arcsine square root transformation' in the IUPAC Compendium of Chemical Terminology, 5th ed.; International Union of Pure and Applied Chemistry; 2025. Online version 5.0.0, 2025. 10.1351/goldbook.14445 • License: The IUPAC Gold Book is licensed under Creative Commons Attribution-ShareAlike CC BY-SA 4.0 International for individual terms. Requests for commercial usage of the compendium should be directed to IUPAC.

  • Peer Review Report
  • 10.7554/elife.78634.sa0
Editor's evaluation: Robust and Efficient Assessment of Potency (REAP) as a quantitative tool for dose-response curve estimation
  • May 9, 2022
  • Philip Boonstra

Finding a new drug which is both safe and efficient is an expensive and time-consuming endeavour. In particular, establishing the ‘dose-effect relationship’ – how beneficial a drug is at different dosages – can be challenging. Predicting this curve requires gathering experimental data by exposing and recording how cells respond to various levels of the drug. However, extreme values are often observed at low and high dosages, potentially introducing errors that are hard to correct in the prediction process. Yet, these extreme observations are sometimes genuine so researchers cannot just ignore them. To improve dose-effect estimation, Zhou, Liu, Fang et al. developed a new general-purpose approach. It uses advanced statistical modelling to account for extremes in lab data. This strategy outperformed other methods when dealing with these observations while also providing higher efficiency in data analysis with more uniform data in experiments. To facilitate implementation, Zhou, Liu, Fang et al. set up a user-friendly tool baptized ‘REAP’; this free online resource allows scientists without advanced statistical experience to harness the new approach and to perform dose-effect analysis more easily and accurately. This could boost research across many different disciplines that examine the effects of chemicals on cells.

  • Discussion
  • Cite Count Icon 13
  • 10.1016/j.ajog.2020.11.017
Vertical transmission of coronavirus disease 2019, a response
  • Nov 20, 2020
  • American Journal of Obstetrics and Gynecology
  • Alexander M Kotlyar + 2 more

Vertical transmission of coronavirus disease 2019, a response

  • Discussion
  • Cite Count Icon 70
  • 10.1016/j.jclinepi.2011.06.016
Logistic regression modeling and the number of events per variable: selection bias dominates
  • Oct 25, 2011
  • Journal of Clinical Epidemiology
  • Ewout W Steyerberg + 2 more

Logistic regression modeling and the number of events per variable: selection bias dominates

  • Research Article
  • Cite Count Icon 3
  • 10.1002/pst.2170
A note on confidence intervals for the restricted mean survival time based on transformations in small sample size.
  • Sep 21, 2021
  • Pharmaceutical Statistics
  • Hiroya Hashimoto + 1 more

Restricted mean survival time (RMST) is one measure now used to summarize time-to-event type data, but it has been pointed out that the distribution of differences in RMST deviates markedly from a normal distribution for controlled clinical trials with small sample sizes. Therefore, we conducted a numerical simulation of the RMST in which the one-sample survival time follows a Weibull distribution, comparing eight different confidence intervals combining two types of variance with four types of variable transformations, including no transformation. The evaluation items were the coverage probability and the above and below error probabilities for the true value. The variance types were based on Greenwood's formula and its Kaplan-Meier correction. The arcsine square root transformation, logit transformation, and complementary log-log transformation were used as the variable transformations. When the sample size was small and the event rate was low, the confidence interval of the untransformed RMST tended to have a small coverage probability and to be overestimated. Variance by Kaplan-Meier correction improved the coverage. The problems of coverage and overestimation were also improved by variable transformations, and in particular, applying the logit transformation and the complementary log-log transformation both resulted in substantial improvements. Our study suggested that it is preferable to construct the confidence intervals of RMST using the logit transformation for variances based on Greenwood's formula in small sample size trials. The SAS code to replicate the analyses is available at https://github.com/HiroyaHashimoto/SAS-Programs.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 143
  • 10.1186/1471-2288-11-77
Logistic random effects regression models: a comparison of statistical packages for binary and ordinal outcomes
  • May 23, 2011
  • BMC Medical Research Methodology
  • Baoyue Li + 3 more

BackgroundLogistic random effects models are a popular tool to analyze multilevel also called hierarchical data with a binary or ordinal outcome. Here, we aim to compare different statistical software implementations of these models.MethodsWe used individual patient data from 8509 patients in 231 centers with moderate and severe Traumatic Brain Injury (TBI) enrolled in eight Randomized Controlled Trials (RCTs) and three observational studies. We fitted logistic random effects regression models with the 5-point Glasgow Outcome Scale (GOS) as outcome, both dichotomized as well as ordinal, with center and/or trial as random effects, and as covariates age, motor score, pupil reactivity or trial. We then compared the implementations of frequentist and Bayesian methods to estimate the fixed and random effects. Frequentist approaches included R (lme4), Stata (GLLAMM), SAS (GLIMMIX and NLMIXED), MLwiN ([R]IGLS) and MIXOR, Bayesian approaches included WinBUGS, MLwiN (MCMC), R package MCMCglmm and SAS experimental procedure MCMC.Three data sets (the full data set and two sub-datasets) were analysed using basically two logistic random effects models with either one random effect for the center or two random effects for center and trial. For the ordinal outcome in the full data set also a proportional odds model with a random center effect was fitted.ResultsThe packages gave similar parameter estimates for both the fixed and random effects and for the binary (and ordinal) models for the main study and when based on a relatively large number of level-1 (patient level) data compared to the number of level-2 (hospital level) data. However, when based on relatively sparse data set, i.e. when the numbers of level-1 and level-2 data units were about the same, the frequentist and Bayesian approaches showed somewhat different results. The software implementations differ considerably in flexibility, computation time, and usability. There are also differences in the availability of additional tools for model evaluation, such as diagnostic plots. The experimental SAS (version 9.2) procedure MCMC appeared to be inefficient.ConclusionsOn relatively large data sets, the different software implementations of logistic random effects regression models produced similar results. Thus, for a large data set there seems to be no explicit preference (of course if there is no preference from a philosophical point of view) for either a frequentist or Bayesian approach (if based on vague priors). The choice for a particular implementation may largely depend on the desired flexibility, and the usability of the package. For small data sets the random effects variances are difficult to estimate. In the frequentist approaches the MLE of this variance was often estimated zero with a standard error that is either zero or could not be determined, while for Bayesian methods the estimates could depend on the chosen "non-informative" prior of the variance parameter. The starting value for the variance parameter may be also critical for the convergence of the Markov chain.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 397
  • 10.7717/peerj.1114
A comparison of observation-level random effect and Beta-Binomial models for modelling overdispersion in Binomial data in ecology & evolution
  • Jul 21, 2015
  • PeerJ
  • Xavier A Harrison

Overdispersion is a common feature of models of biological data, but researchers often fail to model the excess variation driving the overdispersion, resulting in biased parameter estimates and standard errors. Quantifying and modeling overdispersion when it is present is therefore critical for robust biological inference. One means to account for overdispersion is to add an observation-level random effect (OLRE) to a model, where each data point receives a unique level of a random effect that can absorb the extra-parametric variation in the data. Although some studies have investigated the utility of OLRE to model overdispersion in Poisson count data, studies doing so for Binomial proportion data are scarce. Here I use a simulation approach to investigate the ability of both OLRE models and Beta-Binomial models to recover unbiased parameter estimates in mixed effects models of Binomial data under various degrees of overdispersion. In addition, as ecologists often fit random intercept terms to models when the random effect sample size is low (<5 levels), I investigate the performance of both model types under a range of random effect sample sizes when overdispersion is present. Simulation results revealed that the efficacy of OLRE depends on the process that generated the overdispersion; OLRE failed to cope with overdispersion generated from a Beta-Binomial mixture model, leading to biased slope and intercept estimates, but performed well for overdispersion generated by adding random noise to the linear predictor. Comparison of parameter estimates from an OLRE model with those from its corresponding Beta-Binomial model readily identified when OLRE were performing poorly due to disagreement between effect sizes, and this strategy should be employed whenever OLRE are used for Binomial data to assess their reliability. Beta-Binomial models performed well across all contexts, but showed a tendency to underestimate effect sizes when modelling non-Beta-Binomial data. Finally, both OLRE and Beta-Binomial models performed poorly when models contained <5 levels of the random intercept term, especially for estimating variance components, and this effect appeared independent of total sample size. These results suggest that OLRE are a useful tool for modelling overdispersion in Binomial data, but that they do not perform well in all circumstances and researchers should take care to verify the robustness of parameter estimates of OLRE models.

  • Research Article
  • Cite Count Icon 19
  • 10.4103/aja202212
Associations of sex hormone levels with body mass index (BMI) in men: a cross-sectional study using quantile regression analysis.
  • Jan 1, 2023
  • Asian Journal of Andrology
  • Xin Lv + 5 more

Body mass index (BMI) has been increasing globally in recent decades. Previous studies reported that BMI was associated with sex hormone levels, but the results were generated via linear regression or logistic regression, which would lose part of information. Quantile regression analysis can maximize the use of variable information. Our study compared the associations among different regression models. The participants were recruited from the Center of Reproductive Medicine, The First Hospital of Jilin University (Changchun, China) between June 2018 and June 2019. We used linear, logistic, and quantile regression models to calculate the associations between sex hormone levels and BMI. In total, 448 men were included in this study. The average BMI was 25.7 (standard deviation [s.d.]: 3.7) kg m-2; 29.7% (n = 133) of the participants were normal weight, 45.3% (n = 203) of the participants were overweight, and 23.4% (n = 105) of the participants were obese. The levels of testosterone and estradiol significantly differed among BMI groups (all P < 0.05). In linear regression and logistic regression, BMI was associated with testosterone and estradiol levels (both P < 0.05). In quantile regression, BMI was negatively associated with testosterone levels in all quantiles after adjustment for age (all P < 0.05). BMI was positively associated with estradiol levels in most quantiles (≤80th) after adjustment for age (all P < 0.05). Our study suggested that BMI was one of the influencing factors of testosterone and estradiol. Of note, the quantile regression showed that BMI was associated with estradiol only up to the 80th percentile of estradiol.

  • Research Article
  • Cite Count Icon 1
  • 10.34917/9112192
A Web Based User Interface for Machine Learning Analysis of Health and Education Data
  • Sep 15, 2016
  • Digital Scholarship - UNLV (University of Nevada Reno)
  • Chandani Shrestha

The objective of this thesis is to develop a user friendly web application that will be used to analyse data sets using various machine learning algorithms. The application design follows human computer interaction design guidelines and principles to make a user friendly interface [Shn03]. It uses Linear Regression, Logistic Regression, Backpropagation machine learning algorithms for prediction. This application is built using Java, Play framework, Bootstrap and IntelliJ IDE. Java is used in the backend to create a model that maps the input and output data based on any of the above given learning algorithms while Play Framework and Bootstrap are used to display content in frontend. Play framework is used because it is based on web-friendly architecture. As a result it uses predictable, minimal resources (CPU, memory, threads) for highly scalable applications. It is also developer friendly where changes can be made in the code and hitting the refresh button in browser will update the interface. Bootstrap is used to style the web application and it adds responsiveness to the interface with added feature of cross-browser compatible designs. As a result, the website is responsive and fits the screen size of computer. Using this web application users can predict features, category of the entity in the data sets. User needs to submit data set where each row in the data set must represent attributes of the entity. Once data is submitted the application builds a model using user selected machine learning algorithm logistic regression, linear regression or backpropagation. After the model is developed in second stage of the application user can submit attributes of the entity whose category needs to predicted. The predicted category will be displayed on screen in third stage of the application. The interface of the application shows its current active stage. These models are built using 80% of submitted dataset and remaining 20% is used to test the accuracy of the application. In this thesis, prediction accuracy of each algorithm is tested using UCI breast cancer data sets. When tested on breast cancer data with 10 attributes both Logistic Regression and Backpropagation gave 98.5% accuracy. And when tested on breast cancer data with 31 attributes Logistic Regression gave 92.85% accuracy and Backpropagation gave 94.64%.

  • Research Article
  • Cite Count Icon 26
  • 10.1093/ps/76.2.392
Evaluation of logistic versus linear regression models for predicting pulmonary hypertension syndrome (ascites) using cold exposure or pulmonary artery clamp models in broilers
  • Feb 1, 1997
  • Poultry Science
  • Yk Kirby + 3 more

Evaluation of logistic versus linear regression models for predicting pulmonary hypertension syndrome (ascites) using cold exposure or pulmonary artery clamp models in broilers

  • Research Article
  • 10.1186/s12874-025-02563-9
Multi-group global tests for restricted mean survival time and restricted mean time lost: a variable transformation approach
  • May 3, 2025
  • BMC Medical Research Methodology
  • Shuyu Chen + 5 more

BackgroundRestricted mean survival time (RMST) quantifies survival benefits in single-endpoint analysis, while restricted mean time lost (RMTL) measures event-related time loss in competing risks settings. Both provide clinically intuitive interpretations of treatment effects without relying on proportional hazards assumptions or parametric distributions. While existing RMST/RMTL methods focus primarily on two-group comparisons, multi-arm trials are common in practice. However, asymptotic approaches for these metrics suffer from inflated type I error in small samples, limiting their reliability.MethodsWe propose a global test framework using variable transformation methods (e.g., log, clog-log, arcsine square root, logit), which is applicable to multi-group comparisons of RMST and extends to RMTL in the presence of competing risks. Monte-Carlo simulations were conducted to evaluate type I error and power under various scenarios, and two illustrative examples were provided.ResultsSimulations demonstrated that transformed RMST and RMTL global tests effectively controlled type I error across small samples and high censoring rates, while improving power compared to untransformed methods. For single-endpoint analysis, the RMST arcsine square root transformation is recommended. In competing risks settings, RMTL logit transformation is preferred when the event of interest occurs more frequently than competing events, whereas clog-log transformation performs better when competing events dominate.ConclusionsThe proposed transformation-based global tests offer researchers a flexible, assumption-free tool to compare treatment effects across multiple groups with enhanced reliability and interpretability. Additionally, an R package "compRM" was developed to implement the proposed methods.

  • Research Article
  • 10.12691/ajams-9-2-1
Two-Stage Artificial Neural Network Regression Modelling for Wheezing Risk Factors Among Children - A Case Study of Gatundu Hospital, Kenya
  • Mar 30, 2021
  • American Journal of Applied Mathematics and Statistics
  • Thomas Mageto + 1 more

In Kenya wheezing that leads to asthma development in most cases remain under-diagnosed and under-treated. Currently there is no public supported wheezing and asthma care programmes to optimize care for patients with asthma which greatly compounds diagnosis and treatment of the disease. The aim of this study is therefore to consider and analyse the covariates of childhood wheezing among children below 10 years of age in Kenya, a case study of Gatundu hospital in order to improve the provision of wheezing and asthma care services in medical facilities. The possible risk factors in the study are selected from three major groups of demographic, socioeconomic and geographical location factors related to childhood wheezing. The longitudinal secondary data obtained from Gatundu hospital in Kenya were collected and a total of 584 complete cases were recorded. The predictor variables considered in the study include age of children in months, gender, exclusive breastfeeding, exposure to tobacco smoking, difficult living conditions, residence, atopy, maternal age and preterm births. Due to the binary nature of response variable in which data is recorded as presence or absence of wheezing, the risk factors were modelled using multiple logistic regression and Artificial Neural Network Models. Simple random samples of sizes n = 385 without replacement were selected and p-values at 5% level of significance for the variables were recorded. In multiple logistic regression, the five variables identified as possible risk factors for modelling with p-value less than or equal to 0.05 were selected that includes age of children, exclusive breastfeeding, exposure to tobacco smoking, difficult living conditions and residence that recorded p-values of 0.0151, 0.0000, 0.0071, 0.0274 and 0.0410. The best multiple logistic linear regression model selected was based on Akaike Information Criterion (AIC) criterion that recorded null deviance, residual deviance and AIC of 502.44, 179.57 and 191.57 respectively. The precision and accuracy of the multiple logistic regression model were recorded as 89.2% and 93.3% respectively. The Artificial Neural Network was considered for modelling as well, the model with one-hidden layer with four neurons in the hidden layer recorded precision of 97.1% and accuracy of 39.4% while the rest of the models with one hidden layer recorded precision and accuracy of 0.0% and 65.1% respectively. The Artificial Neural Network model with two-hidden layers were also considered and the Network with one neuron in both layers was selected as better performing model with precision and accuracy of 88.2% and 93.3%. The developed two-stage logistic Artificial Neural Network was found to have better performance compared to multiple linear logistic regression and Artificial Neural Networks since it recorded precision and accuracy of 97.1% and 99.0% respectively and hence recommended for consideration in modelling the risk factor of wheezing among children in Kenya.

  • Research Article
  • Cite Count Icon 145
  • 10.1128/aem.67.5.2129-2135.2001
Comparison of logistic regression and linear regression in modeling percentage data.
  • May 1, 2001
  • Applied and Environmental Microbiology
  • Lihui Zhao + 2 more

Percentage is widely used to describe different results in food microbiology, e.g., probability of microbial growth, percent inactivated, and percent of positive samples. Four sets of percentage data, percent-growth-positive, germination extent, probability for one cell to grow, and maximum fraction of positive tubes, were obtained from our own experiments and the literature. These data were modeled using linear and logistic regression. Five methods were used to compare the goodness of fit of the two models: percentage of predictions closer to observations, range of the differences (predicted value minus observed value), deviation of the model, linear regression between the observed and predicted values, and bias and accuracy factors. Logistic regression was a better predictor of at least 78% of the observations in all four data sets. In all cases, the deviation of logistic models was much smaller. The linear correlation between observations and logistic predictions was always stronger. Validation (accomplished using part of one data set) also demonstrated that the logistic model was more accurate in predicting new data points. Bias and accuracy factors were found to be less informative when evaluating models developed for percentage data, since neither of these indices can compare predictions at zero. Model simplification for the logistic model was demonstrated with one data set. The simplified model was as powerful in making predictions as the full linear model, and it also gave clearer insight in determining the key experimental factors.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant