DNA language model GROVER learns sequence context in the human genome

Melissa Sanabria,Jonas Hirsch,Pierre M Joubert,Anna R Poetsch

doi:10.1038/s42256-024-00872-0

Abstract

Deep-learning models that learn a sense of language on DNA have achieved a high level of performance on genome biological tasks. Genome sequences follow rules similar to natural language but are distinct in the absence of a concept of words. We established byte-pair encoding on the human genome and trained a foundation language model called GROVER (Genome Rules Obtained Via Extracted Representations) with the vocabulary selected via a custom task, next-k-mer prediction. The defined dictionary of tokens in the human genome carries best the information content for GROVER. Analysing learned representations, we observed that trained token embeddings primarily encode information related to frequency, sequence content and length. Some tokens are primarily localized in repeats, whereas the majority widely distribute over the genome. GROVER also learns context and lexical ambiguity. Average trained embeddings of genomic regions relate to functional genomics annotation and thus indicate learning of these structures purely from the contextual relationships of tokens. This highlights the extent of information content encoded by the sequence that can be grasped by GROVER. On fine-tuning tasks addressing genome biology with questions of genome element identification and protein–DNA binding, GROVER exceeds other models’ performance. GROVER learns sequence context, a sense for structure and language rules. Extracting this knowledge can be used to compose a grammar book for the code of life.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Nature Machine Intelligence	Publication Date: Jul 23, 2024
Citations: 2	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

DNA language model GROVER learns sequence context in the human genome

Abstract

Talk to us

Similar Papers

More From: Nature Machine Intelligence

Lead the way for us

Similar Papers

Completion of human Chromosome 21, the Human Genome Project, and Steps towards Understanding Ourselves through Comparative Genomics

Journal of Genetics and Molecular Biology | VOL. 11

01 Sep 2000
Journal of Genetics and Molecular Biology | VOL. 11

Green Day: An Interview with NHGRI Director Eric Green
Eric D Green ... Kevin Davies
GEN Biotechnology | VOL. 2
Eric D Green, et. al.Eric D Green ... Kevin Davies
01 Apr 2023
GEN Biotechnology | VOL. 2

Improved prediction of DNA and RNA binding proteins with deep learning models.
Siwen Wu ... Jun-Tao Guo
Briefings in bioinformatics | VOL. 25
Siwen Wu, et. al.Siwen Wu ... Jun-Tao Guo
23 May 2024
Briefings in bioinformatics | VOL. 25

Pretrained domain-specific language model for natural language processing tasks in the AEC domain
Zhe Zheng ... Jia-Rui Lin
Computers in Industry | VOL. 142
Zhe Zheng, et. al.Zhe Zheng ... Jia-Rui Lin
21 Jun 2022
Computers in Industry | VOL. 142

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

DNA language model GROVER learns sequence context in the human genome

Abstract

Talk to us

Similar Papers

More From: Nature Machine Intelligence