An extended clinical EEG dataset with 15,300 automatically labelled recordings for pathology decoding

Ann-Kathrin Kiessner,Robin T Schirrmeister,Lukas A.W Gemein,Joschka Boedecker,Tonio Ball

doi:10.1016/j.nicl.2023.103482

Ann-Kathrin Kiessner, Robin T Schirrmeister + Show 3 more

Open Access

https://doi.org/10.1016/j.nicl.2023.103482

Copy DOI

Journal: NeuroImage: Clinical	Publication Date: Jan 1, 2023
Citations: 5	License type: cc-by-nc-nd

Affiliation: University of Freiburg

Abstract

Automated clinical EEG analysis using machine learning (ML) methods is a growing EEG research area. Previous studies on binary EEG pathology decoding have mainly used the Temple University Hospital (TUH) Abnormal EEG Corpus (TUAB) which contains approximately 3,000 manually labelled EEG recordings. To evaluate and eventually even improve the generalisation performance of machine learning methods for EEG pathology, decoding larger, publicly available datasets is required. A number of studies addressed the automatic labelling of large open-source datasets as an approach to create new datasets for EEG pathology decoding, but little is known about the extent to which training on larger, automatically labelled dataset affects decoding performances of established deep neural networks. In this study, we automatically created additional pathology labels for the Temple University Hospital (TUH) EEG Corpus (TUEG) based on the medical reports using a rule-based text classifier. We generated a dataset of 15,300 newly labelled recordings, which we call the TUH Abnormal Expansion EEG Corpus (TUABEX), and which is five times larger than the TUAB. Since the TUABEX contains more pathological (75%) than non-pathological (25%) recordings, we then selected a balanced subset of 8,879 recordings, the TUH Abnormal Expansion Balanced EEG Corpus (TUABEXB). To investigate how training on a larger, automatically labelled dataset affects the decoding performance of deep neural networks, we applied four established deep convolutional neural networks (ConvNets) to the task of pathological versus non-pathological classification and compared the performance of each architecture after training on different datasets. The results show that training on the automatically labelled TUABEXB dataset rather than training on the manually labelled TUAB dataset increases accuracies on TUABEXB and even for TUAB itself for some architectures. We argue that automatically labelling of large open-source datasets can be used to efficiently utilise the massive amount of EEG data stored in clinical archives. We make the proposed TUABEXB available open source and thus offer a new dataset for EEG machine learning research.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An extended clinical EEG dataset with 15,300 automatically labelled recordings for pathology decoding

Abstract

Talk to us

Similar Papers

More From: NeuroImage: Clinical

Lead the way for us

Similar Papers

Reaching the ceiling? Empirical scaling behaviour for deep EEG pathology classification
Ann-Kathrin Kiessner ... Tonio Ball
Computers in Biology and Medicine | VOL. 178
Ann-Kathrin Kiessner, et. al.Ann-Kathrin Kiessner ... Tonio Ball
07 Jun 2024
Computers in Biology and Medicine | VOL. 178

Machine-learning-based diagnostics of EEG pathology
Lukas A.W Gemein ... Tonio Ball
NeuroImage | VOL. 220
Lukas A.W Gemein, et. al.Lukas A.W Gemein ... Tonio Ball
10 Jun 2020
NeuroImage | VOL. 220

Automatic Report-Based Labelling of Clinical EEGs for Classifier Training
D Western ... L Canham
-
D Western, et. al.D Western ... L Canham
04 Dec 2021
04 Dec 2021

Artificial intelligence in interdisciplinary life science and drug discovery research.
Jürgen Bajorath
Future science OA | VOL. 8
Jürgen BajorathJürgen Bajorath
08 Mar 2022
Future science OA | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An extended clinical EEG dataset with 15,300 automatically labelled recordings for pathology decoding

Abstract

Talk to us

Similar Papers

More From: NeuroImage: Clinical