The molecular entities in linked data dataset.

Dominik Tomaszuk,Łukasz Szeremeta

doi:10.1016/j.dib.2020.105757

Abstract

The Molecular Entities in Linked Data (MEiLD) dataset comprises data of distinct atoms, molecules, ions, ion pairs, radicals, radical ions, and others that can be identifiable as separately distinguishable chemical entities. The dataset is provided in a JSON-LD format and was generated by the SDFEater, a tool that allows parsing atoms, bonds, and other molecule data. MEiLD contains 349,960 of ‘small’ chemical entities. Our dataset is based on the SDF files and is enriched with additional ontologies and line notation data. As a basis, the Molecular Entities in Linked Data dataset uses the Resource Description Framework (RDF) data model. Saving the data in such a model allows preserving the semantic relations, like hierarchical and associative, between them. To describe chemical molecules, vocabularies such as Chemical Vocabulary for Molecular Entities (CVME) and Simple Knowledge Organization System (SKOS) are used. The dataset can be beneficial, among others, for people concerned with research and development tools for cheminformatics and bioinformatics. In this paper, we describe various methods of access to our dataset. In addition to the MEiLD dataset, we publish the Shapes Constraint Language (SHACL) schema of our dataset and the CVME ontology. The data is available in Mendeley Data.

Highlights

The Molecular Entities in Linked Data (MEiLD) dataset comprises data of distinct atoms, molecules, ions, ion pairs, radicals, radical ions, and others that can be identifiable as separately distinguishable chemical entities
Computer science (Information Systems) Semantic Web, Linked Data Graph Document data was acquired by fetching available public domain documents and generated by a software
In Chemical Vocabulary for Molecular Entities (CVME), molecular entities are modeled as instances of the class cvme:MolecularEntity, which is a subclass of skos:Concept

Summary

Data accessibility

Computer science (Information Systems) Semantic Web, Linked Data Graph Document data was acquired by fetching available public domain documents and generated by a software. The presented dataset of molecular entities is useful because it includes a classification, whereby the relationships between molecular entities and their parents and/or children are described. The provided dataset is useful, because all chemicals in the dataset contain a subsumption relationship, meaning that all of the molecular entries are available to semantic reasoning tools that harness the classification hierarchy. The dataset may be beneficial for the users of information services and systems, along with those who use them through query or inference operations. Resources can be described in collaboration with other datasets and linked to data contributed by other communities

Data description

Dataset preparation

Data sources

Access method 1: path-based access

Access method 2

Access method 3

Access method 4

Declaration of Competing Interest

Full Text

Published Version (Free)

View/Download pdf

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Data in Brief	Publication Date: May 27, 2020
Citations: 2	License type: cc-by

R Discovery Prime

The molecular entities in linked data dataset.

Abstract

Highlights

Summary

Published Version (Free)

Talk to us

Similar Papers

More From: Data in Brief

Lead the way for us

Similar Papers

ISO 21526 Conform Metadata Editor for FAIR Unicode SKOS Thesauri.
Mark R Stöhr ... Raphael W Majeed
Studies in health technology and informatics | VOL. 278
Mark R Stöhr, et. al.Mark R Stöhr ... Raphael W Majeed
24 May 2021
ISO 21526 Conform Metadata Editor for FAIR Unicode SKOS Thesauri.
Mark R Stöhr ... Raphael W Majeed

PAConto: RDF Representation of PACDB Data and Ontology of Infectious Diseases Known to Be Related to Glycan Binding
Elena Solovieva ... Hisashi Narimatsu
-
Elena Solovieva, et. al.Elena Solovieva ... Hisashi Narimatsu
07 Dec 2016
07 Dec 2016

No Pain No Gain: Standards mapping in Latimer Core development
Matt Woodburn ... Sharon Grant
Biodiversity Information Science and Standards | VOL. 7
Matt Woodburn, et. al.Matt Woodburn ... Sharon Grant
21 Sep 2023
Biodiversity Information Science and Standards | VOL. 7

SKOS: Simple Knowledge Organisation for the Web
Alistair Miles ... José R Pérez-Agüera
Cataloging & Classification Quarterly | VOL. 43
Alistair Miles, et. al.Alistair Miles ... José R Pérez-Agüera
30 Apr 2007
Cataloging & Classification Quarterly | VOL. 43

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

The molecular entities in linked data dataset.

Abstract

Highlights

Summary

Published Version (Free)

Talk to us

Similar Papers

More From: Data in Brief