A community effort to optimize sequence-based deep learning models of gene regulation.

Abdul Muntakim Rafi,Daria Nogina,Dmitry Penzar,Dohoon Lee,Danyeong Lee,Nayeon Kim,Sangyeup Kim,Dohyeon Kim,Yeojin Shin,Il-Youp Kwak,Georgy Meshcheryakov,Andrey Lando,Arsenii Zinkevich,Byeong-Chan Kim,Juhyun Lee,Taein Kang,Eeshit Dhaval Vaishnav,Payman Yadollahpour,Random Promoter Dream Challenge Consortium ,Susanne Bornelöv,Fredrik Svensson,Maria-Anna Trapotsi,Duc Tran,Tin Nguyen,Xinming Tu,Wuwei Zhang,Wei Qiu,Rohan Ghotra,Yiyang Yu,Ethan Labelson,Aayush Prakash,Ashwin Narayanan,Peter Koo,Xiaoting Chen,David T Jones,Michele Tinti,Yuanfang Guan,Maolin Ding,Ken Chen,Yuedong Yang,Ke Ding,Gunjan Dixit,Jiayu Wen,Zhihan Zhou,Pratik Dutta,Rekha Sathian,Pallavi Surana,Yanrong Ji,Han Liu,Ramana V Davuluri,Yu Hiratsuka,Mao Takatsu,Tsai-Min Chen,Chih-Han Huang,Hsuan-Kai Wang,Edward S C Shih,Sz-Hau Chen,Chih-Hsun Wu,Jhih-Yu Chen,Kuei-Lin Huang,Ibrahim Alsaggaf,Patrick Greaves,Carl Barton,Cen Wan,Nicholas Abad,Cindy Körner,Lars Feuerbach,Benedikt Brors,Yichao Li,Sebastian Röner,Pyaree Mohan Dash,Max Schubach,Onuralp Soylemez,Andreas Møller,Gabija Kavaliauskaite,Jesper Madsen,Zhixiu Lu,Owen Queen,Ashley Babjac,Scott Emrich,Konstantinos Kardamiliotis,Konstantinos Kyriakidis,Andigoni Malousi,Ashok Palaniappan,Krishnakant Gupta,Prasanna Kumar S,Jake Bradford,Dimitri Perrin,Robert Salomone,Carl Schmitz,Chen Jiaxing,Wang Jingzhe,Yang Aiwei,Sun Kim,Jake Albrecht,Aviv Regev,Wuming Gong,Ivan V Kulakovskiy,Pablo Meyer,Carl G De Boer

doi:10.1038/s41587-024-02414-w

Abstract

A systematic evaluation of how model architectures and training strategies impact genomics model performance is needed. To address this gap, we held a DREAM Challenge where competitors trained models on a dataset of millions of random promoter DNA sequences and corresponding expression levels, experimentally determined in yeast. For a robust evaluation of the models, we designed a comprehensive suite of benchmarks encompassing various sequence types. All top-performing models used neural networks but diverged in architectures and training strategies. To dissect how architectural and training choices impact performance, we developed the Prix Fixe framework to divide models into modular building blocks. We tested all possible combinations for the top three models, further improving their performance. The DREAM Challenge models not only achieved state-of-the-art results on our comprehensive yeast dataset but also consistently surpassed existing benchmarks on Drosophila and human genomic datasets, demonstrating the progress that can be driven by gold-standard genomics datasets.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A community effort to optimize sequence-based deep learning models of gene regulation.

Abstract

Talk to us

Similar Papers

More From: Nature biotechnology

Lead the way for us

Journal: Nature biotechnology	Publication Date: Oct 11, 2024
License type: CC BY 4.0

Similar Papers

Modular Filter Building Block for Modular full-SiC AC-DC Converters by an Arrangement of Coupled Inductors
Sungjae Ohn ... Rolando Burgos
-
Sungjae Ohn, et. al.Sungjae Ohn ... Rolando Burgos
11 Oct 2020
11 Oct 2020

Liuhua 11-1 Development-Subsea Production System Overview
Johnce E Hall ... Wang Zhu Sheng
-
Johnce E Hall, et. al.Johnce E Hall ... Wang Zhu Sheng
06 May 1996
06 May 1996

Modular Opto-Mechanical Design of Free-Space Optical Interconnect System for Massively Parallel Processing
Kenneth Fasanella ... David T Neilson
-
Kenneth Fasanella, et. al.Kenneth Fasanella ... David T Neilson
01 Jan 1997
01 Jan 1997

A New Transportation System for the Automatic Handling of Bulk Materials and Suplies in Offshore and Underwater Environments
Arthur W Radford ... Frederick A.F Cooke
-
Arthur W Radford, et. al.Arthur W Radford ... Frederick A.F Cooke
17 May 1969
17 May 1969

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A community effort to optimize sequence-based deep learning models of gene regulation.

Abstract

Talk to us

Similar Papers

More From: Nature biotechnology