Conformal efficiency as a metric for comparative model assessment befitting federated learning

Wouter Heyndrickx,Adam Arany,Jaak Simm,Anastasia Pentina,Noé Sturm,Lina Humbeck,Lewis Mervin,Adam Zalewski,Martijn Oldenhof,Peter Schmidtke,Lukas Friedrich,Regis Loeb,Arina Afanasyeva,Ansgar Schuffenhauer,Yves Moreau,Hugo Ceulemans

doi:10.1016/j.ailsci.2023.100070

Wouter Heyndrickx, Adam Arany + Show 14 more

Open Access

https://doi.org/10.1016/j.ailsci.2023.100070

Copy DOI

Journal: Artificial Intelligence in the Life Sciences	Publication Date: Apr 1, 2023
Citations: 2	License type: cc-by-nc-nd

Abstract

In a drug discovery setting, pharmaceutical companies own substantial but confidential datasets. The MELLODDY project developed a privacy-preserving federated machine learning solution and deployed it at an unprecedented scale. Each partner built models for their own private assays that benefitted from a shared representation. Established predictive performance metrics such as AUC ROC or AUC PR are constrained to unseen labeled chemical space and cannot gage performance gains in unlabeled chemical space. Federated learning indirectly extends labeled space, but in a privacy-preserving context, a partner cannot use this label extension for performance assessment. Metrics that estimate uncertainty on a prediction can be calculated even where no label is known. Practically, the chemical space covered with predictions above an uncertainty threshold, reflects the applicability domain of a model. After establishing a link to established performance metrics, we propose the efficiency from the conformal prediction framework (‘conformal efficiency’) as a proxy to the applicability domain size. A documented extension of the applicability domain would qualify as a tangible benefit from federated learning. In interim assessments, MELLODDY partners reported a median increase in conformal efficiency of the federated over the single-partner model of 5.5% (with increases up to 9.7%). Subject to distributional conditions, that efficiency increase can be directly interpreted as the expected increase in conformal i.e. low uncertainty predictions. In conclusion, we present the first indication that privacy-preserving federated machine learning across massive drug-discovery datasets from ten pharma partners indeed extends the applicability domain of property prediction models.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Conformal efficiency as a metric for comparative model assessment befitting federated learning

Abstract

Talk to us

Similar Papers

More From: Artificial Intelligence in the Life Sciences

Lead the way for us

Similar Papers

Expanding the Applicability Domain of Machine Learning Model for Advancements in Electrochemical Material Discovery
Kajjana Boonpalit ... Jiramet Kinchagawat
ChemElectroChem | VOL. 11
Kajjana Boonpalit, et. al.Kajjana Boonpalit ... Jiramet Kinchagawat
09 Feb 2024
ChemElectroChem | VOL. 11

Efficiency of different measures for defining the applicability domain of classification models
Waldemar Klingspohn ... Miriam Mathea
Journal of Cheminformatics | VOL. 9
Waldemar Klingspohn, et. al.Waldemar Klingspohn ... Miriam Mathea
03 Aug 2017
Journal of Cheminformatics | VOL. 9

Pedestrian Behavior Prediction for Automated Driving: Requirements, Metrics, and Relevant Features
Michael Herman ... Lutz Burkle
IEEE Transactions on Intelligent Transportation Systems | VOL. 23
Michael Herman, et. al.Michael Herman ... Lutz Burkle
01 Sep 2022
IEEE Transactions on Intelligent Transportation Systems | VOL. 23

Development of new methods for the (Q)SAR applicability domain assessment : using structural information in a statistical study of the errors in prediction

-

01 Jan 2015
01 Jan 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Conformal efficiency as a metric for comparative model assessment befitting federated learning

Abstract

Talk to us

Similar Papers

More From: Artificial Intelligence in the Life Sciences