Protein feature engineering framework for AMPylation site prediction

Hardik Prabhu,Hrushikesh Bhosale,Aamod Sane,Renu Dhadwal,Vigneshwar Ramakrishnan,Jayaraman Valadi

doi:10.1038/s41598-024-58450-8

Abstract

AMPylation is a biologically significant yet understudied post-translational modification where an adenosine monophosphate (AMP) group is added to Tyrosine and Threonine residues primarily. While recent work has illuminated the prevalence and functional impacts of AMPylation, experimental identification of AMPylation sites remains challenging. Computational prediction techniques provide a faster alternative approach. The predictive performance of machine learning models is highly dependent on the features used to represent the raw amino acid sequences. In this work, we introduce a novel feature extraction pipeline to encode the key properties relevant to AMPylation site prediction. We utilize a recently published dataset of curated AMPylation sites to develop our feature generation framework. We demonstrate the utility of our extracted features by training various machine learning classifiers, on various numerical representations of the raw sequences extracted with the help of our framework. Tenfold cross-validation is used to evaluate the model’s capability to distinguish between AMPylated and non-AMPylated sites. The top-performing set of features extracted achieved MCC score of 0.58, Accuracy of 0.8, AUC-ROC of 0.85 and F1 score of 0.73. Further, we elucidate the behaviour of the model on the set of features consisting of monogram and bigram counts for various representations using SHapley Additive exPlanations.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Scientific Reports	Publication Date: Apr 15, 2024
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Protein feature engineering framework for AMPylation site prediction

Abstract

Talk to us

Similar Papers

More From: Scientific Reports

Lead the way for us

Similar Papers

Machine learning prediction model of major adverse outcomes after pediatric congenital heart surgery: a retrospective cohort study.
Chaoyang Tong ... Xinwei Du
International Journal of Surgery | VOL. 110
Chaoyang Tong, et. al.Chaoyang Tong ... Xinwei Du
01 Apr 2024
International Journal of Surgery | VOL. 110

Applying machine learning to the pharmacokinetic modeling of cyclosporine in adult renal transplant recipients: a multi-method comparison.
Junjun Mao ... Mingkang Zhong
Frontiers in Pharmacology | VOL. 13
Junjun Mao, et. al.Junjun Mao ... Mingkang Zhong
24 Oct 2022
Frontiers in Pharmacology | VOL. 13

Does Artificial Intelligence Outperform Natural Intelligence in Interpreting Musculoskeletal Radiological Studies? A Systematic Review.
Olivier Q Groot ... Michiel E R Bongers
Clinical Orthopaedics & Related Research | VOL. 478
Olivier Q Groot, et. al.Olivier Q Groot ... Michiel E R Bongers
30 Jul 2020
Clinical Orthopaedics & Related Research | VOL. 478

Computed Tomography Image Analysis on COVID-19 Cases using Machine Learning Approaches
Jasmine Wang Thye Wei ... Ong Kok Haur
Journal of Advanced Research in Applied Sciences and Engineering Technology | VOL. 32
Jasmine Wang Thye Wei, et. al. Jasmine Wang Thye Wei ... Ong Kok Haur
07 Sep 2023
Journal of Advanced Research in Applied Sciences and Engineering Technology | VOL. 32

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Protein feature engineering framework for AMPylation site prediction

Abstract

Talk to us

Similar Papers

More From: Scientific Reports