Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Collinearity: a review of methods to deal with it and a simulation study evaluating their performance

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Collinearity refers to the non independence of predictor variables, usually in a regression‐type analysis. It is a common feature of any descriptive ecological data set and can be a problem for parameter estimation because it inflates the variance of regression parameters and hence potentially leads to the wrong identification of relevant predictors in a statistical model. Collinearity is a severe problem when a model is trained on data from one region or time, and predicted to another with a different or unknown structure of collinearity. To demonstrate the reach of the problem of collinearity in ecology, we show how relationships among predictors differ between biomes, change over spatial scales and through time. Across disciplines, different approaches to addressing collinearity problems have been developed, ranging from clustering of predictors, threshold‐based pre‐selection, through latent variable methods, to shrinkage and regularisation. Using simulated data with five predictor‐response relationships of increasing complexity and eight levels of collinearity we compared ways to address collinearity with standard multiple regression and machine‐learning approaches. We assessed the performance of each approach by testing its impact on prediction to new data. In the extreme, we tested whether the methods were able to identify the true underlying relationship in a training dataset with strong collinearity by evaluating its performance on a test dataset without any collinearity. We found that methods specifically designed for collinearity, such as latent variable methods and tree based models, did not outperform the traditional GLM and threshold‐based pre‐selection. Our results highlight the value of GLM in combination with penalised methods (particularly ridge) and threshold‐based pre‐selection when omitted variables are considered in the final interpretation. However, all approaches tested yielded degraded predictions under change in collinearity structure and the ‘folk lore’‐thresholds of correlation coefficients between predictor variables of |r| >0.7 was an appropriate indicator for when collinearity begins to severely distort model estimation and subsequent prediction. The use of ecological understanding of the system in pre‐analysis variable selection and the choice of the least sensitive statistical approaches reduce the problems of collinearity, but cannot ultimately solve them.

Similar Papers
  • Research Article
  • Cite Count Icon 5
  • 10.1016/j.ijrobp.2018.02.152
Method for Automatic Selection of Parameters in Normal Tissue Complication Probability Modeling
  • Mar 6, 2018
  • International Journal of Radiation Oncology*Biology*Physics
  • Damianos Christophides + 4 more

Method for Automatic Selection of Parameters in Normal Tissue Complication Probability Modeling

  • Research Article
  • Cite Count Icon 2
  • 10.1093/eurpub/ckaa165.267
Use of artificial intelligence to estimate population health indicators in France
  • Sep 1, 2020
  • European Journal of Public Health
  • R Haneef + 6 more

Background The use of artificial intelligence is increasing to estimate and predict health outcomes from large data sets. The main objectives were to develop two algorithms using machine learning techniques to identify new cases of diabetes (case study I) and to classify type 1 and type 2 (case study II) in France. Methods We selected the training data set from a cohort study linked with French national Health database (i.e., SNDS). Two final datasets were used to achieve each objective. A supervised machine learning method including eight following steps was developed: the selection of the data set, case definition, coding and standardization of variables, split data into training and test data sets, variable selection, training, validation and selection of the model. We planned to apply the trained models on the SNDS to estimate the incidence of diabetes and the prevalence of type 1/2 diabetes. Results For the case study I, 23/3468 and for case study II, 14/3481 SNDS variables were selected based on an optimal balance between variance explained and using the ReliefExp algorithm. We trained four models using different classification algorithms on the training data set. The Linear Discriminant Analysis model performed best in both case studies. The models were assessed on the test datasets and achieved a specificity of 67% and a sensitivity of 62% in case study I, and a specificity of 97 % and sensitivity of 100% in case study II. The case study II model was applied to the SNDS and estimated the prevalence of type 1 diabetes in 2016 in France of 0.3% and for type 2, 4.4%. The case study model I was not applied to the SNDS. Conclusions The case study II model to estimate the prevalence of type 1/2 diabetes has good performance and will be used in routine surveillance. The case study I model to identify new cases of diabetes showed a poor performance due to missing necessary information on determinants of diabetes and will need to be improved for further research.

  • Research Article
  • Cite Count Icon 4
  • 10.1016/j.ecolind.2024.112517
Forest above-ground biomass estimation based on strongly collinear variables derived from airborne laser scanning data
  • Aug 30, 2024
  • Ecological Indicators
  • Xiaofang Zhang + 9 more

Forest above-ground biomass estimation based on strongly collinear variables derived from airborne laser scanning data

  • Research Article
  • 10.22067/jsw.v31i3.57184
مدل سازی آماری شوری خاک در پهنه های گسترده
  • Aug 23, 2017
  • پژوهشهای آب و خاک
  • یوسف هاشمی نژاد + 2 more

شور شدن خاک‌ها در جهان به گونه‌ای روزافزون روبه گسترش است و درنتیجه تولید محصولات کشاورزی در مواجهه با این تنش کاهش می‌یابد. سیاست‌گذاران و تصمیم‌سازان در راستای برنامه‌ریزی برای تطبیق با تغییرات اقلیمی و افزایش نیاز به غذا نیازمند پایش کمی مستمر شوری خاک می-باشند. شاخص‌های طیفی حاصل از سنجنده‌های ماهواره‌ای و یا سنجنده‌های نزدیک به سطح زمین به‌طور روزافزونی برای پایش شوری خاک مورداستفاده قرار می‌گیرند به‌نحوی‌که تا کنون تعداد زیادی شاخص برای پایش شوری خاک معرفی شده‌اند. برای مدل‌سازی و سنجش اعتبار مدل حاصله روش‌های رگرسیونی مختلفی مورداستفاده قرار گرفته که مهم‌ترین آن‌ها رگرسیون خطی چندگانه (شامل رگرسیون گام‌به‌گام، انتخاب رو به جلو و حذف رو به عقب) و رگرسیون حداقل مربعات جزئی است. در این پژوهش به‌منظور ارزیابی این دو روش در مدل‌سازی تغییرات شوری خاک از اندازه-گیری‌های آزمایشگاهی و الکترومغناطیسی شوری خاک مربوط به 97 نقطه در سال 1392 و 225 نقطه در سال 1393 در بخشی از دشت سبزوار- داورزن به مساحت حدود 50 هزار هکتار استفاده شد. تعداد 23 شاخص طیفی از تصاویر ماهواره لندست 8 مربوط به تاریخ‌های نمونه‌برداری استخراج و به همراه مدل رقومی ارتفاع به‌عنوان متغیر مستقل مورداستفاده قرار گرفت. روش‌های مختلف رگرسیون خطی چندمتغیره با استفاده از داده‌های سال اول به‌عنوان آموزش و سال دوم به‌عنوان آزمون و بالعکس هرچند ضریب تبیین بین حدود 22 تا 88 درصد ایجاد کرد، ولی این همبستگی در دسته اعتبار سنجی از 29 درصد تجاوز نکرد. به علت وجود هم‌راستایی خطی چندگانه در بین متغیرهای مستقل روش رگرسیون خطی چندگانه برای تمام متغیر‌ها قابل کاربرد نبود. حذف متغیرهای دارای هم‌راستایی خطی، تبدیل لگاریتمی و تصادفی کردن کل داده‌ها در دو دسته آموزش و آزمون، ضریب رگرسیون مدل و اعتبار آن را به‌طور قابل قبولی افزایش داد. استفاده از رگرسیون حداقل مربعات جزئی با استفاده از داده‌های اصلی و تبدیل لگاریتمی شده سال اول و دوم به‌عنوان آموزش و آزمون و بالعکس نیز در دسته آموزش ضریب تبیین بین 39 تا 85 درصد ایجاد کرد، ولی از برآورد در دسته آزمون ناتوان بود. تصادفی کردن داده‌ها و تقسیم مجدد آن‌ها به دو دسته آموزش و آزمون موجب ارتقای چشمگیر ضریب تعیین در دسته اعتبارسنجی شد. تکرار عملیات تصادفی کردن نشان داد که روش از ثبات لازم برای برآورد ضرایب متغیرها برخوردار است.

  • Research Article
  • Cite Count Icon 7
  • 10.1016/j.jpeds.2021.12.058
Development and Validation of a Prediction Model for Infant Fat Mass
  • Dec 28, 2021
  • The Journal of Pediatrics
  • Jasmine F Plows + 9 more

Development and Validation of a Prediction Model for Infant Fat Mass

  • Research Article
  • Cite Count Icon 104
  • 10.1016/s0016-2361(01)00121-1
Neural network prediction of cetane number and density of diesel fuel from its chemical composition determined by LC and GC–MS
  • Aug 10, 2001
  • Fuel
  • Hong Yang + 5 more

Neural network prediction of cetane number and density of diesel fuel from its chemical composition determined by LC and GC–MS

  • Research Article
  • Cite Count Icon 24
  • 10.1016/j.csbj.2022.03.025
Automated coronary artery calcium scoring using nested U-Net and focal loss
  • Jan 1, 2022
  • Computational and Structural Biotechnology Journal
  • Jia-Sheng Hong + 7 more

Automated coronary artery calcium scoring using nested U-Net and focal loss

  • Research Article
  • Cite Count Icon 1
  • 10.1007/s00330-025-11597-y
Multivariable model to predict breast cancer in non-mass enhancement lesions: a study on contrast-enhanced mammography.
  • Apr 17, 2025
  • European radiology
  • Bei Hua + 7 more

To explore morphology and enhancement features of malignant non-mass enhancement (NME) lesions in contrast-enhanced mammography (CEM), and to develop a multivariable model that can accurately predict the probability of malignancy in NME lesions. A total of 162 patients with 206 NME lesions were enrolled. The ratio of 7:3 was randomly divided into a training data set and a test data set. Differences between benign and malignant NME diseases were compared using statistical analysis in the training data set. A logistic regression analysis was used to develop a multivariable model for predicting the probability of malignancy in the training data set. The predictive value of the model was assessed by calculating the area under the curve (AUC) in both training and test data sets. The incidence of malignancy was higher in cases with malignant microcalcification (32.35%), segmental and linear distribution (55.88%), clumped and clustered ring enhancement pattern (70.59%), and Type III curve (64.71%) (all p < 0.002). The sensitivity, specificity, and AUC of the multivariable model in the training data set and the test data set were 79.41-80.77%, 94.44-97.37%, and 0.920-0.946, respectively. When combining microcalcification and enhancement features, the multivariable model for CEM demonstrated acceptable sensitivity and high specificity in predicting malignant NME lesions. Question CEM has gained momentum as an innovative and clinically useful method, but it has not been identified for the discrimination efficacy of NME lesions. Findings The multivariable model of CEM can improve the diagnostic efficiency of breast malignancy NME lesions, with acceptable sensitivity and high specificity. Clinical relevance CEM is an innovative advancement in breast imaging technology. This multivariable model of CEM integrates factors such as microcalcifications, enhancement morphological distribution, internal enhancement patterns, and time-signal intensity curves, thereby enabling accurate diagnosis of NME lesions.

  • Preprint Article
  • 10.21203/rs.3.rs-6097732/v1
Combined model of radiomics and clinical features for predicting prognosis of term neonatal hypoxic-ischemic encephalopathy after one year: an exploratory study
  • Jul 14, 2025
  • Research Square
  • Jing Tang + 5 more

Purpose The purpose of this study was to establish a combined model based on T1WI, T2WI, FLAIR images and clinical parameters to predict the prognosis of hypoxic-ischemic encephalopathy (HIE) in full-term newborns. Methods Based on the results of cognitive scores and motor function scores at 12 months post-birth, the patients were classified into two groups: Group B for those with good prognosis (n = 84) and Group W for those with poor prognosis (n = 96). A total of 180 patients were retrospectively evaluated and assigned to either the training data set (n = 126) or the testing data set (n = 54). The clinical characteristics of both groups were compared first. Then, a clinical model, a radiomics model, and a combined model were developed. Finally, we evaluated the performance of the three constructed models using the receiver operating characteristic (ROC) curve and the area under the curve (AUC). Results The Apgar scores at 1 minute, 5 minutes, and 10 minutes were all higher in Group B compared to Group W, with P &lt; 0.05 indicating a statistically significant difference. The clinical model showed that the Apgar score at 10 minutes was the most effective factor, with an AUC of 0.857 in the training data set and an AUC of 0.737 in the testing data set. For the radiomics model, 9 radiomics features were found to be significantly related to predicting the prognosis of HIE, with AUCs of 0.916 and 0.770 in the training and testing data sets, respectively. For the combined model, 7 radiomics features, Apgar scores at 5 minutes, and Apgar scores at 10 minutes were independent predictors for predicting the prognosis of HIE, with AUCs of 0.952 and 0.823 in the training and testing data sets, respectively. The combined model demonstrates better performance than both the clinical and radiomics models. Conclusions The combined model, which incorporates MR-based radiomics signatures, and clinical factors, is effective in predicting the prognosis of HIE.

  • Conference Article
  • Cite Count Icon 13
  • 10.1109/ebbt.2018.8391441
Classification Of a bank data set on various data mining platforms
  • Apr 1, 2018
  • Muhammet Sinan Basarslan + 1 more

The process of extracting meaningful rules from big and complex data is called data mining. Data mining has an increasing popularity in every field today. Data units are established in customer-oriented industries such as marketing, finance and telecommunication to work on the customer churn and acquisition, in particular. Among the data mining methods, classification algorithms are used in studies conducted for customer acquisition to predict the potential customers of the company in question in the related industry. In this study, bank marketing data set in UCI Machine Learning Data Set was used by creating models with the same classification algorithms in different data mining programs. Accuracy, precision and f- measure criteria were used to test performances of the classification models. When creating the classification models, the test and training data sets were randomly divided by the holdout method to evaluate the performance of the data set. The data set was divided into training and test data sets with the 60-40%, 75­25% and 80-20% separation ratios. Data mining programs used for these processes are the R, Knime, RapidMiner and WEKA. And, classification algorithms commonly used in these platforms are the k-nearest neighbor (k-nn), Naive Bayes, and C4.5 decision tree.

  • Front Matter
  • Cite Count Icon 45
  • 10.1088/0967-3334/33/9/e01
Signal quality in cardiorespiratory monitoring
  • Aug 17, 2012
  • Physiological Measurement
  • Gari D Clifford + 1 more

This focus issue of Physiological Measurement follows the 38th Annual International Computing in Cardiology (CinC) Conference, hosted in Hangzhou, China in September 2011 by Zhejiang University. Each year, the NIH-sponsored PhysioNet resource (http://physionet.org/) runs an open competition lasting several months, aimed at encouraging the development of solutions to an unsolved or poorly solved problem in biomedicine, in most cases making use of relevant clinical and experimental data provided freely by PhysioNet. Participants in these annual challenges discuss their diverse approaches to the Challenge problems during dedicated scientific sessions at CinC. The topics of these PhysioNet/CinC Challenges range from physiologic signal processing and analysis to forecasting and modelling clinically important events and processes.

  • Research Article
  • Cite Count Icon 16
  • 10.1016/j.joms.2020.02.007
A Validated Model to Predict Postoperative Symptom Severity After Mandibular Third Molar Removal
  • Feb 12, 2020
  • Journal of Oral and Maxillofacial Surgery
  • Feng Qiao + 5 more

A Validated Model to Predict Postoperative Symptom Severity After Mandibular Third Molar Removal

  • Research Article
  • Cite Count Icon 92
  • 10.1542/peds.2004-2099
Prediction of Death for Extremely Low Birth Weight Neonates
  • Dec 1, 2005
  • Pediatrics
  • Namasivayam Ambalavanan + 9 more

To compare multiple logistic regression and neural network models in predicting death for extremely low birth weight neonates at 5 time points with cumulative data sets, as follows: scenario A, limited prenatal data; scenario B, scenario A plus additional prenatal data; scenario C, scenario B plus data from the first 5 minutes after birth; scenario D, scenario C plus data from the first 24 hours after birth; scenario E, scenario D plus data from the first 1 week after birth. Data for all infants with birth weights of 401 to 1000 g who were born between January 1998 and April 2003 in 19 National Institute of Child Health and Human Development Neonatal Research Network centers were used (n = 8608). Twenty-eight variables were selected for analysis (3 for scenario A, 15 for scenario B, 20 for scenario C, 25 for scenario D, and 28 for scenario E) from those collected routinely. Data sets censored for prior death or missing data were created for each scenario and divided randomly into training (70%) and test (30%) data sets. Logistic regression and neural network models for predicting subsequent death were created with training data sets and evaluated with test data sets. The predictive abilities of the models were evaluated with the area under the curve of the receiver operating characteristic curves. The data sets for scenarios A, B, and C were similar, and prediction was best with scenario C (area under the curve: 0.85 for regression; 0.84 for neural networks), compared with scenarios A and B. The logistic regression and neural network models performed similarly well for scenarios A, B, D, and E, but the regression model was superior for scenario C. Prediction of death is limited even with sophisticated statistical methods such as logistic regression and nonlinear modeling techniques such as neural networks. The difficulty of predicting death should be acknowledged in discussions with families and caregivers about decisions regarding initiation or continuation of care.

  • Research Article
  • Cite Count Icon 21
  • 10.1002/jcb.28028
The discovery of a novel eight-mRNA-lncRNA signature predicting survival of hepatocellular carcinoma patients.
  • Nov 28, 2018
  • Journal of Cellular Biochemistry
  • Ye‐Min Shi + 5 more

Increasing evidence indicates that the expressions of messenger RNAs (mRNAs) and long non-coding RNAs (lncRNAs) undergo a frequent and aberrant change in carcinogenesis and cancer development. But some research was carried out on mRNA-lncRNA signatures for prediction of hepatocellular carcinoma (HCC) prognosis. We aimed to establish an mRNA-lncRNA signature to improve the ability to predict HCC patients' survival. The subjects from the cancer genome atlas (TCGA) data set were randomly divided into two parts: training data set (n = 246) and testing data set (n = 124). Using computational methods, we selected eight gene signatures (five mRNAs and three lncRNAs) to generate the risk score model, which were significantly correlated with overall survival of patients withHCC in both training and testing data set. The signature had the ability to classify the patients in training data set into a high-risk group and low-risk group with significantly different overall survival (hazard ratio = 4.157, 95% confidence interval = 2.648-6.526, P < 0.001). The prognostic value was further validated in testing data set and the entire data set. Further analysis revealed that this signature was independent of tumor stage. In addition, Gene Set Enrichment Analysis suggested that high risk score group was associated with cell proliferation and division related pathways. Finally, we developed a well-performed nomogram integrating the prognostic signature and other clinical information to predict 3- and 5-year overall survival. In conclusion, the prognostic mRNAs and lncRNAs identified in our study indicate their potential role in HCC biogenesis. The risk score model based on the mRNA-lncRNA may be an efficient classification tool to evaluate the prognosis of patients' withHCC.

  • PDF Download Icon
  • Research Article
  • 10.1088/1757-899x/302/1/012036
Learning Data Set Influence on Identification Accuracy of Gas Turbine Neural Network Model
  • Jan 1, 2018
  • IOP Conference Series: Materials Science and Engineering
  • A V Kuznetsov + 1 more

There are many gas turbine engine identification researches via dynamic neural network models. It should minimize errors between model and real object during identification process. Questions about training data set processing of neural networks are usually missed. This article presents a study about influence of data set type on gas turbine neural network model accuracy. The identification object is thermodynamic model of micro gas turbine engine. The thermodynamic model input signal is the fuel consumption and output signal is the engine rotor rotation frequency. Four types input signals was used for creating training and testing data sets of dynamic neural network models – step, fast, slow and mixed. Four dynamic neural networks were created based on these types of training data sets. Each neural network was tested via four types test data sets. In the result 16 transition processes from four neural networks and four test data sets from analogous solving results of thermodynamic model were compared. The errors comparison was made between all neural network errors in each test data set. In the comparison result it was shown error value ranges of each test data set. It is shown that error values ranges is small therefore the influence of data set types on identification accuracy is low.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant