Collinearity: a review of methods to deal with it and a simulation study evaluating their performance
Collinearity refers to the non independence of predictor variables, usually in a regression‐type analysis. It is a common feature of any descriptive ecological data set and can be a problem for parameter estimation because it inflates the variance of regression parameters and hence potentially leads to the wrong identification of relevant predictors in a statistical model. Collinearity is a severe problem when a model is trained on data from one region or time, and predicted to another with a different or unknown structure of collinearity. To demonstrate the reach of the problem of collinearity in ecology, we show how relationships among predictors differ between biomes, change over spatial scales and through time. Across disciplines, different approaches to addressing collinearity problems have been developed, ranging from clustering of predictors, threshold‐based pre‐selection, through latent variable methods, to shrinkage and regularisation. Using simulated data with five predictor‐response relationships of increasing complexity and eight levels of collinearity we compared ways to address collinearity with standard multiple regression and machine‐learning approaches. We assessed the performance of each approach by testing its impact on prediction to new data. In the extreme, we tested whether the methods were able to identify the true underlying relationship in a training dataset with strong collinearity by evaluating its performance on a test dataset without any collinearity. We found that methods specifically designed for collinearity, such as latent variable methods and tree based models, did not outperform the traditional GLM and threshold‐based pre‐selection. Our results highlight the value of GLM in combination with penalised methods (particularly ridge) and threshold‐based pre‐selection when omitted variables are considered in the final interpretation. However, all approaches tested yielded degraded predictions under change in collinearity structure and the ‘folk lore’‐thresholds of correlation coefficients between predictor variables of |r| >0.7 was an appropriate indicator for when collinearity begins to severely distort model estimation and subsequent prediction. The use of ecological understanding of the system in pre‐analysis variable selection and the choice of the least sensitive statistical approaches reduce the problems of collinearity, but cannot ultimately solve them.
- Research Article
5
- 10.1016/j.ijrobp.2018.02.152
- Mar 6, 2018
- International Journal of Radiation Oncology*Biology*Physics
Method for Automatic Selection of Parameters in Normal Tissue Complication Probability Modeling
- Research Article
2
- 10.1093/eurpub/ckaa165.267
- Sep 1, 2020
- European Journal of Public Health
Background The use of artificial intelligence is increasing to estimate and predict health outcomes from large data sets. The main objectives were to develop two algorithms using machine learning techniques to identify new cases of diabetes (case study I) and to classify type 1 and type 2 (case study II) in France. Methods We selected the training data set from a cohort study linked with French national Health database (i.e., SNDS). Two final datasets were used to achieve each objective. A supervised machine learning method including eight following steps was developed: the selection of the data set, case definition, coding and standardization of variables, split data into training and test data sets, variable selection, training, validation and selection of the model. We planned to apply the trained models on the SNDS to estimate the incidence of diabetes and the prevalence of type 1/2 diabetes. Results For the case study I, 23/3468 and for case study II, 14/3481 SNDS variables were selected based on an optimal balance between variance explained and using the ReliefExp algorithm. We trained four models using different classification algorithms on the training data set. The Linear Discriminant Analysis model performed best in both case studies. The models were assessed on the test datasets and achieved a specificity of 67% and a sensitivity of 62% in case study I, and a specificity of 97 % and sensitivity of 100% in case study II. The case study II model was applied to the SNDS and estimated the prevalence of type 1 diabetes in 2016 in France of 0.3% and for type 2, 4.4%. The case study model I was not applied to the SNDS. Conclusions The case study II model to estimate the prevalence of type 1/2 diabetes has good performance and will be used in routine surveillance. The case study I model to identify new cases of diabetes showed a poor performance due to missing necessary information on determinants of diabetes and will need to be improved for further research.
- Research Article
4
- 10.1016/j.ecolind.2024.112517
- Aug 30, 2024
- Ecological Indicators
Forest above-ground biomass estimation based on strongly collinear variables derived from airborne laser scanning data
- Research Article
- 10.22067/jsw.v31i3.57184
- Aug 23, 2017
- پژوهشهای آب و خاک
شور شدن خاکها در جهان به گونهای روزافزون روبه گسترش است و درنتیجه تولید محصولات کشاورزی در مواجهه با این تنش کاهش مییابد. سیاستگذاران و تصمیمسازان در راستای برنامهریزی برای تطبیق با تغییرات اقلیمی و افزایش نیاز به غذا نیازمند پایش کمی مستمر شوری خاک می-باشند. شاخصهای طیفی حاصل از سنجندههای ماهوارهای و یا سنجندههای نزدیک به سطح زمین بهطور روزافزونی برای پایش شوری خاک مورداستفاده قرار میگیرند بهنحویکه تا کنون تعداد زیادی شاخص برای پایش شوری خاک معرفی شدهاند. برای مدلسازی و سنجش اعتبار مدل حاصله روشهای رگرسیونی مختلفی مورداستفاده قرار گرفته که مهمترین آنها رگرسیون خطی چندگانه (شامل رگرسیون گامبهگام، انتخاب رو به جلو و حذف رو به عقب) و رگرسیون حداقل مربعات جزئی است. در این پژوهش بهمنظور ارزیابی این دو روش در مدلسازی تغییرات شوری خاک از اندازه-گیریهای آزمایشگاهی و الکترومغناطیسی شوری خاک مربوط به 97 نقطه در سال 1392 و 225 نقطه در سال 1393 در بخشی از دشت سبزوار- داورزن به مساحت حدود 50 هزار هکتار استفاده شد. تعداد 23 شاخص طیفی از تصاویر ماهواره لندست 8 مربوط به تاریخهای نمونهبرداری استخراج و به همراه مدل رقومی ارتفاع بهعنوان متغیر مستقل مورداستفاده قرار گرفت. روشهای مختلف رگرسیون خطی چندمتغیره با استفاده از دادههای سال اول بهعنوان آموزش و سال دوم بهعنوان آزمون و بالعکس هرچند ضریب تبیین بین حدود 22 تا 88 درصد ایجاد کرد، ولی این همبستگی در دسته اعتبار سنجی از 29 درصد تجاوز نکرد. به علت وجود همراستایی خطی چندگانه در بین متغیرهای مستقل روش رگرسیون خطی چندگانه برای تمام متغیرها قابل کاربرد نبود. حذف متغیرهای دارای همراستایی خطی، تبدیل لگاریتمی و تصادفی کردن کل دادهها در دو دسته آموزش و آزمون، ضریب رگرسیون مدل و اعتبار آن را بهطور قابل قبولی افزایش داد. استفاده از رگرسیون حداقل مربعات جزئی با استفاده از دادههای اصلی و تبدیل لگاریتمی شده سال اول و دوم بهعنوان آموزش و آزمون و بالعکس نیز در دسته آموزش ضریب تبیین بین 39 تا 85 درصد ایجاد کرد، ولی از برآورد در دسته آزمون ناتوان بود. تصادفی کردن دادهها و تقسیم مجدد آنها به دو دسته آموزش و آزمون موجب ارتقای چشمگیر ضریب تعیین در دسته اعتبارسنجی شد. تکرار عملیات تصادفی کردن نشان داد که روش از ثبات لازم برای برآورد ضرایب متغیرها برخوردار است.
- Research Article
7
- 10.1016/j.jpeds.2021.12.058
- Dec 28, 2021
- The Journal of Pediatrics
Development and Validation of a Prediction Model for Infant Fat Mass
- Research Article
104
- 10.1016/s0016-2361(01)00121-1
- Aug 10, 2001
- Fuel
Neural network prediction of cetane number and density of diesel fuel from its chemical composition determined by LC and GC–MS
- Research Article
24
- 10.1016/j.csbj.2022.03.025
- Jan 1, 2022
- Computational and Structural Biotechnology Journal
Automated coronary artery calcium scoring using nested U-Net and focal loss
- Research Article
1
- 10.1007/s00330-025-11597-y
- Apr 17, 2025
- European radiology
To explore morphology and enhancement features of malignant non-mass enhancement (NME) lesions in contrast-enhanced mammography (CEM), and to develop a multivariable model that can accurately predict the probability of malignancy in NME lesions. A total of 162 patients with 206 NME lesions were enrolled. The ratio of 7:3 was randomly divided into a training data set and a test data set. Differences between benign and malignant NME diseases were compared using statistical analysis in the training data set. A logistic regression analysis was used to develop a multivariable model for predicting the probability of malignancy in the training data set. The predictive value of the model was assessed by calculating the area under the curve (AUC) in both training and test data sets. The incidence of malignancy was higher in cases with malignant microcalcification (32.35%), segmental and linear distribution (55.88%), clumped and clustered ring enhancement pattern (70.59%), and Type III curve (64.71%) (all p < 0.002). The sensitivity, specificity, and AUC of the multivariable model in the training data set and the test data set were 79.41-80.77%, 94.44-97.37%, and 0.920-0.946, respectively. When combining microcalcification and enhancement features, the multivariable model for CEM demonstrated acceptable sensitivity and high specificity in predicting malignant NME lesions. Question CEM has gained momentum as an innovative and clinically useful method, but it has not been identified for the discrimination efficacy of NME lesions. Findings The multivariable model of CEM can improve the diagnostic efficiency of breast malignancy NME lesions, with acceptable sensitivity and high specificity. Clinical relevance CEM is an innovative advancement in breast imaging technology. This multivariable model of CEM integrates factors such as microcalcifications, enhancement morphological distribution, internal enhancement patterns, and time-signal intensity curves, thereby enabling accurate diagnosis of NME lesions.
- Preprint Article
- 10.21203/rs.3.rs-6097732/v1
- Jul 14, 2025
- Research Square
Purpose The purpose of this study was to establish a combined model based on T1WI, T2WI, FLAIR images and clinical parameters to predict the prognosis of hypoxic-ischemic encephalopathy (HIE) in full-term newborns. Methods Based on the results of cognitive scores and motor function scores at 12 months post-birth, the patients were classified into two groups: Group B for those with good prognosis (n = 84) and Group W for those with poor prognosis (n = 96). A total of 180 patients were retrospectively evaluated and assigned to either the training data set (n = 126) or the testing data set (n = 54). The clinical characteristics of both groups were compared first. Then, a clinical model, a radiomics model, and a combined model were developed. Finally, we evaluated the performance of the three constructed models using the receiver operating characteristic (ROC) curve and the area under the curve (AUC). Results The Apgar scores at 1 minute, 5 minutes, and 10 minutes were all higher in Group B compared to Group W, with P < 0.05 indicating a statistically significant difference. The clinical model showed that the Apgar score at 10 minutes was the most effective factor, with an AUC of 0.857 in the training data set and an AUC of 0.737 in the testing data set. For the radiomics model, 9 radiomics features were found to be significantly related to predicting the prognosis of HIE, with AUCs of 0.916 and 0.770 in the training and testing data sets, respectively. For the combined model, 7 radiomics features, Apgar scores at 5 minutes, and Apgar scores at 10 minutes were independent predictors for predicting the prognosis of HIE, with AUCs of 0.952 and 0.823 in the training and testing data sets, respectively. The combined model demonstrates better performance than both the clinical and radiomics models. Conclusions The combined model, which incorporates MR-based radiomics signatures, and clinical factors, is effective in predicting the prognosis of HIE.
- Conference Article
13
- 10.1109/ebbt.2018.8391441
- Apr 1, 2018
The process of extracting meaningful rules from big and complex data is called data mining. Data mining has an increasing popularity in every field today. Data units are established in customer-oriented industries such as marketing, finance and telecommunication to work on the customer churn and acquisition, in particular. Among the data mining methods, classification algorithms are used in studies conducted for customer acquisition to predict the potential customers of the company in question in the related industry. In this study, bank marketing data set in UCI Machine Learning Data Set was used by creating models with the same classification algorithms in different data mining programs. Accuracy, precision and f- measure criteria were used to test performances of the classification models. When creating the classification models, the test and training data sets were randomly divided by the holdout method to evaluate the performance of the data set. The data set was divided into training and test data sets with the 60-40%, 7525% and 80-20% separation ratios. Data mining programs used for these processes are the R, Knime, RapidMiner and WEKA. And, classification algorithms commonly used in these platforms are the k-nearest neighbor (k-nn), Naive Bayes, and C4.5 decision tree.
- Front Matter
45
- 10.1088/0967-3334/33/9/e01
- Aug 17, 2012
- Physiological Measurement
This focus issue of Physiological Measurement follows the 38th Annual International Computing in Cardiology (CinC) Conference, hosted in Hangzhou, China in September 2011 by Zhejiang University. Each year, the NIH-sponsored PhysioNet resource (http://physionet.org/) runs an open competition lasting several months, aimed at encouraging the development of solutions to an unsolved or poorly solved problem in biomedicine, in most cases making use of relevant clinical and experimental data provided freely by PhysioNet. Participants in these annual challenges discuss their diverse approaches to the Challenge problems during dedicated scientific sessions at CinC. The topics of these PhysioNet/CinC Challenges range from physiologic signal processing and analysis to forecasting and modelling clinically important events and processes.
- Research Article
16
- 10.1016/j.joms.2020.02.007
- Feb 12, 2020
- Journal of Oral and Maxillofacial Surgery
A Validated Model to Predict Postoperative Symptom Severity After Mandibular Third Molar Removal
- Research Article
92
- 10.1542/peds.2004-2099
- Dec 1, 2005
- Pediatrics
To compare multiple logistic regression and neural network models in predicting death for extremely low birth weight neonates at 5 time points with cumulative data sets, as follows: scenario A, limited prenatal data; scenario B, scenario A plus additional prenatal data; scenario C, scenario B plus data from the first 5 minutes after birth; scenario D, scenario C plus data from the first 24 hours after birth; scenario E, scenario D plus data from the first 1 week after birth. Data for all infants with birth weights of 401 to 1000 g who were born between January 1998 and April 2003 in 19 National Institute of Child Health and Human Development Neonatal Research Network centers were used (n = 8608). Twenty-eight variables were selected for analysis (3 for scenario A, 15 for scenario B, 20 for scenario C, 25 for scenario D, and 28 for scenario E) from those collected routinely. Data sets censored for prior death or missing data were created for each scenario and divided randomly into training (70%) and test (30%) data sets. Logistic regression and neural network models for predicting subsequent death were created with training data sets and evaluated with test data sets. The predictive abilities of the models were evaluated with the area under the curve of the receiver operating characteristic curves. The data sets for scenarios A, B, and C were similar, and prediction was best with scenario C (area under the curve: 0.85 for regression; 0.84 for neural networks), compared with scenarios A and B. The logistic regression and neural network models performed similarly well for scenarios A, B, D, and E, but the regression model was superior for scenario C. Prediction of death is limited even with sophisticated statistical methods such as logistic regression and nonlinear modeling techniques such as neural networks. The difficulty of predicting death should be acknowledged in discussions with families and caregivers about decisions regarding initiation or continuation of care.
- Research Article
21
- 10.1002/jcb.28028
- Nov 28, 2018
- Journal of Cellular Biochemistry
Increasing evidence indicates that the expressions of messenger RNAs (mRNAs) and long non-coding RNAs (lncRNAs) undergo a frequent and aberrant change in carcinogenesis and cancer development. But some research was carried out on mRNA-lncRNA signatures for prediction of hepatocellular carcinoma (HCC) prognosis. We aimed to establish an mRNA-lncRNA signature to improve the ability to predict HCC patients' survival. The subjects from the cancer genome atlas (TCGA) data set were randomly divided into two parts: training data set (n = 246) and testing data set (n = 124). Using computational methods, we selected eight gene signatures (five mRNAs and three lncRNAs) to generate the risk score model, which were significantly correlated with overall survival of patients withHCC in both training and testing data set. The signature had the ability to classify the patients in training data set into a high-risk group and low-risk group with significantly different overall survival (hazard ratio = 4.157, 95% confidence interval = 2.648-6.526, P < 0.001). The prognostic value was further validated in testing data set and the entire data set. Further analysis revealed that this signature was independent of tumor stage. In addition, Gene Set Enrichment Analysis suggested that high risk score group was associated with cell proliferation and division related pathways. Finally, we developed a well-performed nomogram integrating the prognostic signature and other clinical information to predict 3- and 5-year overall survival. In conclusion, the prognostic mRNAs and lncRNAs identified in our study indicate their potential role in HCC biogenesis. The risk score model based on the mRNA-lncRNA may be an efficient classification tool to evaluate the prognosis of patients' withHCC.
- Research Article
- 10.1088/1757-899x/302/1/012036
- Jan 1, 2018
- IOP Conference Series: Materials Science and Engineering
There are many gas turbine engine identification researches via dynamic neural network models. It should minimize errors between model and real object during identification process. Questions about training data set processing of neural networks are usually missed. This article presents a study about influence of data set type on gas turbine neural network model accuracy. The identification object is thermodynamic model of micro gas turbine engine. The thermodynamic model input signal is the fuel consumption and output signal is the engine rotor rotation frequency. Four types input signals was used for creating training and testing data sets of dynamic neural network models – step, fast, slow and mixed. Four dynamic neural networks were created based on these types of training data sets. Each neural network was tested via four types test data sets. In the result 16 transition processes from four neural networks and four test data sets from analogous solving results of thermodynamic model were compared. The errors comparison was made between all neural network errors in each test data set. In the comparison result it was shown error value ranges of each test data set. It is shown that error values ranges is small therefore the influence of data set types on identification accuracy is low.