Leveraging Scheme for Cross-Study Microbiome Machine Learning Prediction and Feature Evaluations.

Kuncheng Song,Yi-Hui Zhou

doi:10.3390/bioengineering10020231

Kuncheng Song, Yi-Hui Zhou

Open Access

https://doi.org/10.3390/bioengineering10020231

Copy DOI

Journal: Bioengineering	Publication Date: Feb 8, 2023
Citations: 2	License type: CC BY 4.0

Affiliation: North Carolina State University

Abstract

The microbiota has proved to be one of the critical factors for many diseases, and researchers have been using microbiome data for disease prediction. However, models trained on one independent microbiome study may not be easily applicable to other independent studies due to the high level of variability in microbiome data. In this study, we developed a method for improving the generalizability and interpretability of machine learning models for predicting three different diseases (colorectal cancer, Crohn's disease, and immunotherapy response) using nine independent microbiome datasets. Our method involves combining a smaller dataset with a larger dataset, and we found that using at least 25% of the target samples in the source data resulted in improved model performance. We determined random forest as our top model and employed feature selection to identify common and important taxa for disease prediction across the different studies. Our results suggest that this leveraging scheme is a promising approach for improving the accuracy and interpretability of machine learning models for predicting diseases based on microbiome data.

Full Text