Classification of metabolites by metabolic pathways concerning terpenoids, phenylpropanoids, and polyketide compounds based on machine learning
Terpenoids, phenylpropanoids, and polyketides are the majority of the secondary metabolites containing carbon, hydrogen, and oxygen. In this work, 19,769 metabolites accumulated in KNApSAcK Core DB were classified into 71 subgroups comprising three major groups (terpenoids, phenylpropanoids, and polyketides) according to scientific literatures. We represented the metabolites as molecular fingerprint including chemical properties, and used those descriptors for classification by random forest model. We found that both training and test metabolites were well classified into the subgroups, with 94.06 %, and 94.23 % accuracy, respectively. Though classification of metabolites based on metabolic pathways is very time-consuming works, machine learnings with molecular fingerprint made it possible to attain the classification. This work will lead a light for systematical and evolutional understanding of diverged secondary metabolites based on secondary metabolic pathways. Data science is an interdisciplinary and applied field that uses techniques and theories drawn from statistics, mathematics, computer science, and information science. Combining these resources data science enables extracting meaningful and practical insights for secondary metabolites.
- Research Article
4
- 10.1016/j.clnu.2025.05.011
- Jul 1, 2025
- Clinical nutrition (Edinburgh, Scotland)
Integration of metabolomics and machine learning for precise management and prevention of cardiometabolic risk in Asians.
- Research Article
22
- 10.1002/jcsm.13045
- Jul 21, 2022
- Journal of cachexia, sarcopenia and muscle
ObjectivesIdiopathic inflammatory myopathies (IIM) are a class of autoimmune diseases with high heterogeneity that can be divided into different subtypes based on clinical manifestations and myositis‐specific autoantibodies (MSAs). However, even in each IIM subtype, the clinical symptoms and prognoses of patients are very different. Thus, the identification of more potential biomarkers associated with IIM classification, clinical symptoms, and prognosis is urgently needed.MethodsPlasma and urine samples from 79 newly diagnosed IIM patients (mean disease duration 4 months) and 52 normal control (NC) samples were analysed by high‐performance liquid chromatography of quadrupole time‐of‐flight mass spectrometry (HPLC‐Q‐TOF‐MS)/MS‐based untargeted metabolomics. Orthogonal partial least‐squares discriminate analysis (OPLS‐DA) were performed to measure the significance of metabolites. Pathway enrichment analysis was conducted based on the KEGG human metabolic pathways. Ten machine learning (ML) algorithms [linear support vector machine (SVM), radial basis function SVM, random forest, nearest neighbour, Gaussian processes, decision trees, neural networks, adaptive boosting (AdaBoost), Gaussian naive Bayes and quadratic discriminant analysis] were used to classify each IIM subtype and select the most important metabolites as potential biomarkers.ResultsOPLS‐DA showed a clear separation between NC and IIM subtypes in plasma and urine metabolic profiles. KEGG pathway enrichment analysis revealed multiple unique and shared disturbed metabolic pathways in IIM main [dermatomyositis (DM), anti‐synthetase syndrome (ASS), and immune‐mediated necrotizing myopathy (IMNM)] and MSA‐defined subtypes (anti‐Mi2+, anti‐MDA5+, anti‐TIF1γ+, anti‐Jo1+, anti‐PL7+, anti‐PL12+, anti‐EJ+, and anti‐SRP+), such that fatty acid biosynthesis was significantly altered in both plasma and urine in all main IIM subtypes (enrichment ratio > 1). Random forest and AdaBoost performed best in classifying each IIM subtype among the 10 ML models. Using the feature selection methods in ML models, we identified 9 plasma and 10 urine metabolites that contributed most to separate IIM main subtypes and MSA‐defined subtypes, such as plasma creatine (fold change = 3.344, P = 0.024) in IMNM subtype and urine tiglylcarnitine (fold change = 0.351, P = 0.037) in anti‐EJ+ ASS subtype. Sixteen common metabolites were found in both the plasma and urine samples of IIM subtypes. Among them, some were correlated with clinical features, such as plasma hypogeic acid (r = −0.416, P = 0.005) and urine malonyl carnitine (r = −0.374, P = 0.042), which were negatively correlated with the prevalence of interstitial lung disease.ConclusionsIn both plasma and urine samples, IIM main and MSA‐defined subtypes have specific metabolic signatures and pathways. This study provides useful clues for understanding the molecular mechanisms, searching potential diagnosis biomarkers and therapeutic targets for IIM.
- Research Article
11
- 10.1016/j.clim.2024.110235
- May 6, 2024
- Clinical Immunology
Targeted metabolomics combined with machine learning to identify and validate new biomarkers for early SLE diagnosis and disease activity
- Research Article
- 10.1186/s12864-026-12734-7
- Mar 14, 2026
- BMC genomics
Meat quality traits are typically regulated by multiple genes, each contributing a small effect. In this study, to pinpoint candidate genes involved in meat quality traits, we performed transcriptome profiles of porcine longissimus dorsi (LD) muscle and applied machine learning (ML) models to analyze RNA-seq data. We also carried out Gene Set Enrichment Analysis (ssGSEA), Weighted gene co-expression network analysis (WGCNA) and functional validation of putative target genes to better support the biological relevance of our findings. In this study, LD muscle samples were collected from 142 Huoshou Black (HSH) pigs and 191 Anqing Six-end-white (AQLB) pigs. Based on results of the estimated breeding values (EBV) analysis, of meat quality traits, we selected 101 HSH pigs and 99 AQLB pigs for transcriptomic analysis. Using an integrative analytical framework that combined ssGSEA and WGCNA, we identified 197 candidate genes 197 candidate genes. These genes were significantly associated with various metabolic pathways, including fatty-acid elongation and metabolism, amino-acid catabolism, protein turnover, and biosynthetic processes. To further refine the identification of key regulatory genes, we systematically evaluated ten ML models, ultimately selecting XGBoost, Random Forest, and Lasso Regression for subsequent analysis. This approach pinpointed CYSLTR1 and LPCAT2 as the key regulatory genes. To investigate the functional roles of CYSLTR1 and LPCAT2 in intramuscular fat (IMF) deposition, we established a porcine intramuscular adipocyte model via siRNA-mediated knockdown of either CYSLTR1 or LPCAT2. RNA-Seq analysis identified 339 differentially expressed genes (DEGs) in the siLPCAT2 group and 2,376 DEGs in the siCYSLTR1 group relative to the control. Heatmap analysis indicated that genes involved in triacylglycerol (TAG) biosynthesis were upregulated in the siCYSLTR1 group, but downregulated in the siLPCAT2 group. KEGG pathway enrichment analysis further demonstrated that LPCAT2-associated DEGs were predominantly enriched in the MAPK signaling pathway, mTOR signaling pathway, biosynthesis of unsaturated fatty acids, and glycerolipid metabolism—pathways closely linked to cell proliferation, nutrient sensing, and lipid remodeling. In contrast, CYSLTR1-associated DEGs were significantly enriched in lipid metabolism, atherosclerosis, and adipocytokine signaling pathways—processes directly implicated in adipocyte differentiation, lipid storage, and inflammatory crosstalk within adipose tissue. Collectively, these findings elucidate distinct yet complementary regulatory roles for CYSLTR1 and LPCAT2 in intramuscular adipogenesis and provide mechanistic support for targeting these genes to modulate IMF content and improve pork quality traits. In conclusion, CYSLTR1 and LPCAT2 were identified as pivotal regulatory genes governing IMF deposition. Functional enrichment and pathway analyses revealed that both genes exert their effects on IMF accumulation through the coordinated regulation of lipid metabolism–associated pathways—including fatty acid synthesis, triglyceride assembly, and phospholipid remodeling. These findings offer mechanistically grounded evidence and actionable biological insights for improving pork quality traits, particularly marbling and tenderness.
- Research Article
5
- 10.3389/frai.2022.744755
- Jun 10, 2022
- Frontiers in artificial intelligence
The use of machine learning (ML) in life sciences has gained wide interest over the past years, as it speeds up the development of high performing models. Important modeling tools in biology have proven their worth for pathway design, such as mechanistic models and metabolic networks, as they allow better understanding of mechanisms involved in the functioning of organisms. However, little has been done on the use of ML to model metabolic pathways, and the degree of non-linearity associated with them is not clear. Here, we report the construction of different metabolic pathways with several linear and non-linear ML models. Different types of data are used; they lead to the prediction of important biological data, such as pathway flux and final product concentration. A comparison reveals that the data features impact model performance and highlight the effectiveness of non-linear models (e.g., QRF: RMSE = 0.021 nmol·min−1 and R2 = 1 vs. Bayesian GLM: RMSE = 1.379 nmol·min−1 R2 = 0.823). It turns out that the greater the degree of non-linearity of the pathway, the better suited a non-linear model will be. Therefore, a decision-making support for pathway modeling is established. These findings generally support the hypothesis that non-linear aspects predominate within the metabolic pathways. This must be taken into account when devising possible applications of these pathways for the identification of biomarkers of diseases (e.g., infections, cancer, neurodegenerative diseases) or the optimization of industrial production processes.
- Research Article
1
- 10.2196/85654
- Feb 19, 2026
- Journal of medical Internet research
Stroke is a complex, multidimensional disorder influenced by interacting inflammatory, immune, coagulation, endothelial, and metabolic pathways. Single-omics approaches seldom capture this complexity, whereas multiomics techniques provide complementary insights but generate high-dimensional and correlated feature spaces. Machine learning (ML) offers strategies to manage these challenges; however, the predictive accuracy and reproducibility of multiomics-based ML models for stroke remain poorly characterized. This review aimed to conduct a systematic evaluation of ML models using multiomics data for stroke risk stratification and comprehensive patterns in discriminatory performance, integration strategies, and validation and reporting practices to inform future methodological development. We conducted a comprehensive literature search following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 recommendations. Studies published from January 2000 to July 2025 were identified across 9 databases, including PubMed, MEDLINE Ultimate, EMBASE, CINAHL, Web of Science, Scopus, Cochrane CENTRAL, ACM Digital Library, and IEEE Xplore. Eligible studies included adults with ischemic, hemorrhagic, or unspecified stroke as the prediction target; applied at least 2 omics layers; and reported ML performance metrics. Risk of bias was assessed using the Prediction Model Risk of Bias Assessment Tool, while reporting quality was evaluated using Minimum Information for Medical AI Reporting. The primary outcome was the area under the receiver operating characteristic curve. A total of 7 studies (n=40,274) published between 2022 and 2025 fulfilled the inclusion criteria. All studies combined 2 omics layers, most often using middle-level integration with dyads such as metabolomics-proteomics and metabolomics-lipidomics. Supervised ML algorithms across studies included support vector machines, tree-based ensembles, generalized linear models, and deep learning architectures. Three studies reported external validation of the integrated multiomics model, while 1 study conducted only an external assessment of a single marker rather than validation of the integrated model. Three studies reported an assessment of calibration, and clinically prespecified operating points were rarely described. Reported areas under the receiver operating characteristic curve varied by prediction task, ranging from 0.75 to 0.96 for acute diagnosis models and from 0.75 to 0.97 for onset risk prediction models; the highest externally validated performance was achieved by a support vector machine trained on a metabolomics-proteomics dyad in mixed stroke types (ischemic and hemorrhagic). Multiomics ML models showed high apparent discrimination for stroke risk stratification, but current evidence remains methodologically limited. Small sample sizes, heterogeneous designs, and incomplete reporting currently hinder the reproducibility and generalizability of multiomics ML models for stroke risk prediction. To advance the field, future studies should adopt leakage-resistant evaluation frameworks, conduct site-specific external validations, and benchmark against both single-omics and clinical baselines to demonstrate incremental value. Well-designed, transparently reported investigations will be essential to move multiomics ML models from exploratory promise toward clinically actionable tools in precision stroke care.
- Research Article
1
- 10.3389/fmed.2024.1400166
- Sep 20, 2024
- Frontiers in medicine
Sepsis poses a serious threat to individual life and health. Early and accessible diagnosis and targeted treatment are crucial. This study aims to explore the relationship between microbes, metabolic pathways, and blood test indicators in sepsis patients and develop a machine learning model for clinical diagnosis. Blood samples from sepsis patients were sequenced. α-diversity and β-diversity analyses were performed to compare the microbial diversity between the sepsis group and the normal group. Correlation analysis was conducted on microbes, metabolic pathways, and blood test indicators. In addition, a model was developed based on medical records and radiomic features using machine learning algorithms. The results of α-diversity and β-diversity analyses showed that the microbial diversity of sepsis group was significantly higher than that of normal group (p < 0.05). The top 10 microbial abundances in the sepsis and normal groups were Vitis vinifera, Mycobacterium canettii, Solanum pennellii, Ralstonia insidiosa, Ananas comosus, Moraxella osloensis, Escherichia coli, Staphylococcus hominis, Camelina sativa, and Cutibacterium acnes. The enriched metabolic pathways mainly included Protein families: genetic information processing, Translation, Protein families: signaling and cellular processes, and Unclassified: genetic information processing. The correlation analysis revealed a significant positive correlation (p < 0.05) between IL-6 and Membrane transport. Metabolism of other amino acids showed a significant positive correlation (p < 0.05) with Cutibacterium acnes, Ralstonia insidiosa, Moraxella osloensis, and Staphylococcus hominis. Ananas comosus showed a significant positive correlation (p < 0.05) with Poorly characterized and Unclassified: metabolism. Blood test-related indicators showed a significant negative correlation (p < 0.05) with microorganisms. Logistic regression (LR) was used as the optimal model in six machine learning models based on medical records and radiomic features. The nomogram, calibration curves, and AUC values demonstrated that LR performed best for prediction. This study provides insights into the relationship between microbes, metabolic pathways, and blood test indicators in sepsis. The developed machine learning model shows potential for aiding in clinical diagnosis. However, further research is needed to validate and improve the model.
- Research Article
4
- 10.52783/cana.v31.1462
- Sep 2, 2024
- Communications on Applied Nonlinear Analysis
Metagenomics has revolutionized our understanding of microbial communities by enabling the study of genetic material recovered directly from environmental samples. Traditional methods of microbiology often miss the vast majority of microorganisms that are unculturable in laboratory settings. Harnessing the power of machine learning in metagenomics provides an unprecedented opportunity to uncover the diversity and functionality of these invisible microbial worlds. By analyzing large-scale metagenomic datasets, machine learning algorithms can identify patterns and associations that are not easily discernible through conventional analytical techniques, paving the way for new discoveries in microbial ecology and evolution.The integration of machine learning into metagenomics has the potential to enhance the accuracy and speed of taxonomic classification, functional annotation, and the prediction of microbial interactions. Machine learning models can process complex, high-dimensional data, enabling researchers to make more informed predictions about microbial roles in various ecosystems. Additionally, machine learning techniques can aid in identifying novel genes and metabolic pathways that could have significant implications for biotechnology, medicine, and environmental science. These advancements could lead to breakthroughs in areas such as antibiotic resistance, bioremediation, and the development of new bioproducts.As machine learning continues to evolve, its application in metagenomics will likely expand, offering deeper insights into microbial dynamics and their influence on human health and the environment. However, challenges remain, including the need for large, well-curated datasets and the development of models that can handle the complexity and variability of metagenomic data. Despite these challenges, the synergy between machine learning and metagenomics holds great promise for advancing our understanding of the microbial world and unlocking the potential of microbes in various fields.
- Research Article
- 10.1093/intbio/zyag010
- Jan 16, 2026
- Integrative biology : quantitative biosciences from nano to macro
Sepsis, a life-threatening dysregulated host response to infection, involves complex cytokine signaling. Comprehensive bioinformatics analysis of cytokine activity, associated pathways, and immune alterations in sepsis is warranted. Using the sepsis dataset GSE26378 from GEO, we analyzed differential cytokine pathway activity with ssGSEA and identified differentially expressed genes (DEGs). Cytokine-related genes (CRGs) were extracted and overlapped with DEGs. Protein-Protein Interaction (PPI) network analysis and functional enrichment were performed on differentially expressed CRGs. Cytokine activity scores and pathway activities were quantified using Gene Set Variation Analysis (GSVA). Immune cell infiltration was assessed with MCP-counter. Machine learning algorithms (Random Forest, LASSO, SVM) identified diagnostic biomarkers, validated using an independent dataset (GSE26440) and ROC analysis. Cytokine/cytokine receptor pathways were significantly upregulated in sepsis. We identified 617 DEGs and 46 differentially expressed CRGs. Cytokine activity scores were significantly elevated in sepsis and strongly correlated with heightened activity in inflammatory pathways (e.g. TLR, IL-1R, NF-κB, JAK/STAT, hypoxia) and metabolic pathways (e.g. glycolysis, PI3K/AKT/mTOR). Immune analysis showed decreased T cells, NK cells, B cells, and cytotoxic lymphocytes, alongside increased neutrophils and endothelial cells; neutrophil infiltration positively correlated with cytokine scores. Machine learning identified four core genes (C3AR1, XCL1, CSF2RA, IL2RB), consistently dysregulated in sepsis across datasets and demonstrating robust diagnostic accuracy. This integrated bioinformatics study indicates heightened cytokine activity, profound alterations in inflammatory and metabolic pathways, and a dysregulated immune cell landscape in sepsis. The identified hub genes and the four-gene biomarker panel show potential as diagnostic tools, offering insights into sepsis pathophysiology. Insight box This study integrates multi-omics bioinformatics (ssGSEA, GSVA, PPI, immune deconvolution) and machine learning (RF, LASSO, SVM) to dissect sepsis pathophysiology. Innovatively, we quantify cytokine pathway hyperactivity, linking it to inflammatory/metabolic dysregulation (TLR, NF-κB, glycolysis) and immune imbalance. A novel four-gene panel (C3AR1, XCL1, CSF2RA, IL2RB) was identified and validated as a robust diagnostic biomarker, bridging cytokine signaling with clinical utility. The findings provide mechanistic insights into sepsis-driven immune-metabolic crosstalk and offer translational potential for early diagnosis and targeted therapy.
- Research Article
4
- 10.1016/j.ymeth.2024.09.002
- Sep 12, 2024
- Methods
Gluconeogenesis unraveled: A proteomic Odyssey with machine learning
- Research Article
8
- 10.3389/fimmu.2025.1567466
- Jun 20, 2025
- Frontiers in Immunology
ObjectiveMetabolic dysregulation and redox imbalance in immune cells are key drivers of systemic lupus erythematosus (SLE) pathogenesis. This study explores critical oxidative stress (OS) features and their interrelationships in SLE pathogenesis.MethodsThree transcriptomic datasets from the Gene Expression Omnibus (GEO) were analyzed to identify SLE- and OS-associated pathways via Gene Set Variation Analysis (GSVA). Multiple machine learning methods—including deep learning (DL), random forest (RF), XGBoost, support vector machine (SVM), and least absolute shrinkage and selection operator (LASSO)—were deployed to build OS-related gene prediction frameworks. Immune infiltration was assessed using CIBERSORT, and single-cell transcriptomic data from GEO elucidated gene expression patterns in various immune cell subsets. Peripheral blood plasma samples from confirmed SLE patients and healthy controls (HC) were analyzed using liquid chromatography-mass spectrometry (LC-MS) for metabolomics profiling and to evaluate OS and antioxidant stress (AOS) levels. Finally, real-time quantitative PCR (RT-qPCR) was used to validate the expression differences of key genes in peripheral blood mononuclear cells (PBMCs) from SLE patients and HC.ResultsGSVA identified 15 metabolic pathways significantly linked to SLE, seven of which were strongly associated with OS and energy metabolism. LC-MS revealed substantial alterations in serum OS-related metabolites, clearly distinguishing SLE patients from healthy controls. A comprehensive machine learning approach pinpointed 10 OS-related genes; among these, six (ABCB1, AKR1C3, EIF2AK2, IFIH1, NPC1, SCO2) showed robust predictive performance and significant correlations with immune cell subsets. Single-cell analysis confirmed these genes’ expression in diverse immune cell types, consistent with the observed metabolic pathway disruptions. RT-qPCR verified downregulation of ABCB1, AKR1C3, and NPC1 and upregulation of EIF2AK2, IFIH1, and SCO2 in SLE PBMCs. SLE patients exhibited higher OS levels and lower AOS levels. Correlation analysis underscored strong relationships among key genes, OS/AOS levels, and vital metabolites.ConclusionThis multi-omics and machine learning–based investigation uncovered major disruptions in OS-related metabolic pathways and metabolites in SLE, ultimately identifying six key genes with distinct expression patterns across immune cell subsets. Their strong associations with OS/AOS levels and crucial metabolites highlight their diagnostic and therapeutic potential, laying a foundation for early detection and targeted treatment strategies.
- Research Article
29
- 10.1167/iovs.63.1.28
- Jan 21, 2022
- Investigative Ophthalmology & Visual Science
PurposeAdvances in mass spectrometry have provided new insights into the role of metabolomics in the etiology of several diseases. Studies on retinopathy of prematurity (ROP), for example, overlooked the role of metabolic alterations in disease development. We employed comprehensive metabolic profiling and gold-standard metabolic analysis to explore major metabolites and metabolic pathways, which were significantly affected in early stages of pathogenesis toward ROP.MethodsThis was a multicenter, retrospective, matched-pair, case-control study. We collected plasma from 57 ROP cases and 57 strictly matched non-ROP controls. Non-targeted ultra-high-performance liquid chromatography–tandem mass spectroscopy (UPLC-MS/MS) was used to detect the metabolites. Machine learning was employed to reveal the most affected metabolites and pathways in ROP development.ResultsCompared with non-ROP controls, we found a significant metabolic perturbation in the plasma of ROP cases, which featured an increase in the levels of lipids, nucleotides, and carbohydrate metabolites and lower levels of peptides. Machine leaning enabled us to distinguish a cluster of metabolic pathways (glycometabolism, redox homeostasis, lipid metabolism, and arginine pathway) were strongly correlated with the development of ROP. Moreover, the severity of ROP was associated with the levels of creatinine and ribitol; also, overactivity of aerobic glycolysis and lipid metabolism was noted in the metabolic profile of ROP.ConclusionsThe results suggest a strong correlation between metabolic profiling and retinal neovascularization in ROP pathogenesis. These findings provide an insight into the identification of novel metabolic biomarkers for the diagnosis and prevention of ROP, but the clinical significance requires further validation.
- Research Article
- 10.3389/fphar.2026.1768109
- Feb 9, 2026
- Frontiers in pharmacology
Risperidone is one of the most widely prescribed antipsychotics for the management of irritability and associated behavioral symptoms in autism spectrum disorder (ASD), yet clinical response and adverse-effect risk vary widely among individuals. Pharmacogenomic (PGx) research has sought to explain this variability, with accumulating evidence pointing to contributions from metabolic, transporter, and neurotransmitter pathways. In this narrative minireview, we synthesize current findings on PGx factors influencing risperidone outcomes in children and adolescents with ASD. CYP2D6 emerges as the most robust predictor of pharmacokinetics and toxicity, while pharmacodynamic associations involving dopaminergic, serotonergic, and metabolic pathways in genes such as ABCB1, DRD3, HTR2A, HTR2C, and LEP remain inconsistent and largely derived from small cohorts. We also discuss methodological challenges in assessing treatment response, current clinical guidelines, barriers to implementation, and emerging approaches including polygenic models, pharmacoepigenomics, and machine learning. Together, the available evidence points to both the promise and the limitations of PGx in guiding safer and more individualized risperidone therapy in ASD.
- Research Article
3
- 10.1371/journal.pone.0266730
- Aug 16, 2022
- PLOS ONE
To prospectively establish an early diagnosis model of acute colon cancerous bowel obstruction by applying nuclear magnetic resonance hydrogen spectroscopy(1H NMR) technology based metabolomics methods, combined with machine learning. In this study, serum samples of 71 patients with acute bowel obstruction requiring emergency surgery who were admitted to the Emergency Department of Sichuan Provincial People's Hospital from December 2018 to November 2020 were collected within 2 hours after admission, and NMR spectroscopy data was taken after pretreatment. After postoperative pathological confirmation, they were divided into colon cancerous bowel obstruction (CBO) group and adhesive bowel obstruction (ABO) control group. Used MestReNova software to extract the two sets of spectra bins, and used the MetaboAnalyst5.0 website to perform partial least square discrimination (PLS-DA), combining the human metabolome database (HMDB) and the Kyoto Encyclopedia of Genes and Genomes (KEGG) to find possible different Metabolites and related metabolic pathways. 22 patients were classified as CBO group and 30 were classified as ABO control group. Compared with ABO group, the level of Xanthurenic acid, 3-Hydroxyanthranilic acid, Gentisic acid, Salicyluric acid, Ferulic acid, Kynurenic acid, CDP, Mandelic acid, NADPH, FAD, Phenylpyruvate, Allyl isothiocyanate, and Vanillylmandelic acid increased in the CBO group; while the lecel of L-Tryptophan and Bilirubin decreased. There were significant differences between two groups in the tryptophan metabolism, tyrosine metabolism, glutathione metabolism, phenylalanine metabolism and synthesis pathways of phenylalanine, tyrosine and tryptophan (all P<0.05). Tryptophan metabolism pathway had the greatest impact (Impact = 0.19). The early diagnosis model of colon cancerous bowel was established based on the levels of six metabolites: Xanthurenic acid, 3-Hydroxyanthranilic acid, Gentisic acid, Salicylic acid, Ferulic acid and Kynurenic acid (R2 = 0.995, Q2 = 0.931, RMSE = 0.239, AUC = 0.962). This study firstly used serum to determine the difference in metabolome between patients with colon cancerous bowel obstruction and those with adhesive bowel obstruction. The study found that the metabolic information carried by the serum was sufficient to discriminate the two groups of patients and provided the theoretical supporting for the future using of the more convenient sample for the differential diagnosis of patients with colon cancerous bowel obstruction. Quantitative experiments on a large number of samples were still needed in the future.
- Research Article
15
- 10.3389/fbioe.2022.788300
- Jul 7, 2022
- Frontiers in bioengineering and biotechnology
Proteins are some of the most fascinating and challenging molecules in the universe, and they pose a big challenge for artificial intelligence. The implementation of machine learning/AI in protein science gives rise to a world of knowledge adventures in the workhorse of the cell and proteome homeostasis, which are essential for making life possible. This opens up epistemic horizons thanks to a coupling of human tacit–explicit knowledge with machine learning power, the benefits of which are already tangible, such as important advances in protein structure prediction. Moreover, the driving force behind the protein processes of self-organization, adjustment, and fitness requires a space corresponding to gigabytes of life data in its order of magnitude. There are many tasks such as novel protein design, protein folding pathways, and synthetic metabolic routes, as well as protein-aggregation mechanisms, pathogenesis of protein misfolding and disease, and proteostasis networks that are currently unexplored or unrevealed. In this systematic review and biochemical meta-analysis, we aim to contribute to bridging the gap between what we call binomial artificial intelligence (AI) and protein science (PS), a growing research enterprise with exciting and promising biotechnological and biomedical applications. We undertake our task by exploring “the state of the art” in AI and machine learning (ML) applications to protein science in the scientific literature to address some critical research questions in this domain, including What kind of tasks are already explored by ML approaches to protein sciences? What are the most common ML algorithms and databases used? What is the situational diagnostic of the AI–PS inter-field? What do ML processing steps have in common? We also formulate novel questions such as Is it possible to discover what the rules of protein evolution are with the binomial AI–PS? How do protein folding pathways evolve? What are the rules that dictate the folds? What are the minimal nuclear protein structures? How do protein aggregates form and why do they exhibit different toxicities? What are the structural properties of amyloid proteins? How can we design an effective proteostasis network to deal with misfolded proteins? We are a cross-functional group of scientists from several academic disciplines, and we have conducted the systematic review using a variant of the PICO and PRISMA approaches. The search was carried out in four databases (PubMed, Bireme, OVID, and EBSCO Web of Science), resulting in 144 research articles. After three rounds of quality screening, 93 articles were finally selected for further analysis. A summary of our findings is as follows: regarding AI applications, there are mainly four types: 1) genomics, 2) protein structure and function, 3) protein design and evolution, and 4) drug design. In terms of the ML algorithms and databases used, supervised learning was the most common approach (85%). As for the databases used for the ML models, PDB and UniprotKB/Swissprot were the most common ones (21 and 8%, respectively). Moreover, we identified that approximately 63% of the articles organized their results into three steps, which we labeled pre-process, process, and post-process. A few studies combined data from several databases or created their own databases after the pre-process. Our main finding is that, as of today, there are no research road maps serving as guides to address gaps in our knowledge of the AI–PS binomial. All research efforts to collect, integrate multidimensional data features, and then analyze and validate them are, so far, uncoordinated and scattered throughout the scientific literature without a clear epistemic goal or connection between the studies. Therefore, our main contribution to the scientific literature is to offer a road map to help solve problems in drug design, protein structures, design, and function prediction while also presenting the “state of the art” on research in the AI–PS binomial until February 2021. Thus, we pave the way toward future advances in the synthetic redesign of novel proteins and protein networks and artificial metabolic pathways, learning lessons from nature for the welfare of humankind. Many of the novel proteins and metabolic pathways are currently non-existent in nature, nor are they used in the chemical industry or biomedical field.