Perplexity-Based Molecule Ranking and Bias Estimation of Chemical Language Models.

Michael Moret,Gisbert Schneider,Francesca Grisoni,Paul Katzberger

doi:10.1021/acs.jcim.2c00079

Abstract

Chemical language models (CLMs) can be employed to design molecules with desired properties. CLMs generate new chemical structures in the form of textual representations, such as the simplified molecular input line entry system (SMILES) strings. However, the quality of these de novo generated molecules is difficult to assess a priori. In this study, we apply the perplexity metric to determine the degree to which the molecules generated by a CLM match the desired design objectives. This model-intrinsic score allows identifying and ranking the most promising molecular designs based on the probabilities learned by the CLM. Using perplexity to compare “greedy” (beam search) with “explorative” (multinomial sampling) methods for SMILES generation, certain advantages of multinomial sampling become apparent. Additionally, perplexity scoring is performed to identify undesired model biases introduced during model training and allows the development of a new ranking system to remove those undesired biases.

Highlights

Generative deep learning has become a promising method for chemistry and drug discovery.[1−21] Generative models learn the pattern distribution of the input data and generate new data instances based on learned probabilities.[22]
Perplexity has been used to assess the performance of language models in natural language processing.[37−39] For a simplified molecular input line entry system (SMILES) string of length N, the perplexity score can be computed by considering the chemical language models (CLMs) probability of any ith character: N
The information on the overall character probabilities is captured into a single metric, which is normalized by the length of the SMILES string (N)

Summary

Introduction

Generative deep learning has become a promising method for chemistry and drug discovery.[1−21] Generative models learn the pattern distribution of the input data and generate new data instances based on learned probabilities.[22] Among the proposed generative frameworks that have been applied to de novo molecular design,[2−19] chemical language models (CLMs) have gained attention because of their ability to generate focused virtual chemical libraries and bioactive compounds.[20,21,23] CLMs are trained on string representations of molecules, e.g., simplified molecular input line entry system (SMILES) strings (Figure 1a),[24] to iteratively predict the next. Alternative generative approaches have been proposed for de novo design,[13,27−29] benchmarks have not shown these to outperform CLMs.[30,31] A feature of CLMs is their ability to function in low-data regimes,[25,29] i.e., with limited training data (typically in the range of 5−40 molecules).[2,3,25] One of the most widely employed approaches for low-data model training is transfer learning.[20,32] This method leverages previously acquired information on a related task for which more data are available (′′pretraining”) before training the CLM on a more specific limited dataset (′′fine-tuning”).[33]

Objectives

Methods

Results

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Journal of chemical information and modeling	Publication Date: Feb 22, 2022
Citations: 10	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Perplexity-Based Molecule Ranking and Bias Estimation of Chemical Language Models.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of chemical information and modeling

Lead the way for us

Similar Papers

MERMAID: an open source automated hit-to-lead method based on deep reinforcement learning
Daiki Erikawa ... Masakazu Sekijima
Journal of Cheminformatics | VOL. 13
Daiki Erikawa, et. al.Daiki Erikawa ... Masakazu Sekijima
27 Nov 2021
Journal of Cheminformatics | VOL. 13

Can large language models understand molecules?
Shaghayegh Sadeghi ... Alioune Ngom
BMC Bioinformatics | VOL. 25
Shaghayegh Sadeghi, et. al.Shaghayegh Sadeghi ... Alioune Ngom
26 Jun 2024
BMC Bioinformatics | VOL. 25

Simplified molecular input line entry system-based optimal descriptors: QSAR modelling for voltage-gated potassium channel subunit Kv7.2
P Ganga Raju Achary
SAR and QSAR in Environmental Research | VOL. 25
P Ganga Raju AcharyP Ganga Raju Achary
02 Jan 2014
SAR and QSAR in Environmental Research | VOL. 25

QSAR modeling of toxicities of ionic liquids toward Staphylococcus aureus using SMILES and graph invariants
Shahram Lotfi ... Shahin Ahmadi
Structural Chemistry | VOL. 31
Shahram Lotfi, et. al.Shahram Lotfi ... Shahin Ahmadi
09 Jul 2020
Structural Chemistry | VOL. 31

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Perplexity-Based Molecule Ranking and Bias Estimation of Chemical Language Models.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of chemical information and modeling