Abstract

Pre-trained protein language and/or structural models are often fine-tuned on drug development properties (i.e. developability properties) to accelerate drug discovery initiatives. However, these models generally rely on a single structural conformation and/or a single sequence as a molecular representation. We present a physics-based model, whereby 3D conformational ensemble representations are fused by a transformer-based architecture and concatenated to a language representation to predict antibody protein properties. Antibody language ensemble fusion enables the direct infusion of thermodynamic information into latent space and this enhances property prediction by explicitly infusing dynamic molecular behavior that occurs during experimental measurement. We showcase the antibody language ensemble fusion model on two developability properties: hydrophobic interaction chromatography retention time and temperature of aggregation (Tagg). We find that (i) 3D conformational ensembles that are generated from molecular simulation can further improve antibody property prediction for small datasets, (ii) the performance benefit from 3D conformational ensembles matches shallow machine learning methods in the small data regime, and (iii) fine-tuned large protein language models can match smaller antibody-specific language models at predicting antibody properties. AbLEF codebase is available at https://github.com/merck/AbLEF.

Full Text
Paper version not known

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.