Semantic Search in Legacy Biodiversity Literature: Integrating data from different data infrastructures

Adrian Pachzelt,Gerwin Kasperek,Giuseppe Abrami,Andy Lücking,Christine Driller

doi:10.3897/biss.5.74251

Adrian Pachzelt, Gerwin Kasperek + Show 3 more

Open Access

https://doi.org/10.3897/biss.5.74251

Copy DOI

Journal: Biodiversity Information Science and Standards	Publication Date: Sep 10, 2021
Citations: 2	License type: CC BY 4.0

Abstract

Nowadays, obtaining information by entering queries into a web search engine is routine behaviour. With its search portal, the Specialised Information Service Biodiversity Research (BIOfid) adapts the exploration of legacy biodiversity literature and data extraction to current standards (Driller et al. 2020). In this presentation, we introduce the BIOfid search portal and its functionalities in a How-To short guide. To this end, we adapted a knowledge graph representation of our thematic focus of Central European, primarily German language, biodiversity literature of the 19th and 20th centuries. Now, users can search our text-mined corpus containing to date more than 8.700 full-text articles from 68 journals, and particularly focussing on birds, lepidopterans and vascular plants. The texts are automatically preprocessed by the Natural Language Processing provider TextImager (Hemati et al. 2016) and will be linked to various databases such as Wikidata, Wikipedia, the Global Biodiversity Information Facility (GBIF), Encyclopedia of Life (EoL), Geonames, the Integrated Authority File (GND) and WordNet. For data retrieval, users can filter search results and download the article metadata as well as text annotations and database links in JavaScript Object Notation (JSON) format. For example, literature that mentions taxa from certain decades or co-occurrences of species can be searched. Our search engine recognises scientific and vernacular taxon names based on the GBIF Backbone Taxonomy and offers search suggestions to support the user. The semantic network of the BIOfid search portal is also enriched with data from the EoL trait bank, so that trait data can be included in the search queries. Thus, scientists can enhance their own data sets with the search results and feed them into the relevant biodiversity data repositories to sustainably expand the corresponding knowledge graphs with reliable data. Since BIOfid applies standard ontology terms, all data mobilized from literature can be combined with data on natural history collection objects or data from current research projects in order to generate more comprehensive knowledge. Furthermore, taxonomy, ecology and trait ontologies that have been built or extended within this project will be made available through appropriate platforms such as The Open Biological and Biomedical Ontology (OBO) Foundry and the Terminology Service of The German Federation for Biological Data (GFBio).

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Semantic Search in Legacy Biodiversity Literature: Integrating data from different data infrastructures

Abstract

Talk to us

Similar Papers

More From: Biodiversity Information Science and Standards

Lead the way for us

Similar Papers

The Open Biodiversity Knowledge Management (eco-)System: Tools and Services for Extraction, Mobilization, Handling and Re-use of Data from the Published Literature
Lyubomir Penev ... Donat Agosti
Biodiversity Information Science and Standards | VOL. 2
Lyubomir Penev, et. al.Lyubomir Penev ... Donat Agosti
17 May 2018
Biodiversity Information Science and Standards | VOL. 2

The Standards behind the Scenes: Explaining data from the Plazi workflow
Donat Agosti ... Alexandros Ioannidis-Pantopikos
Biodiversity Information Science and Standards | VOL. 4
Donat Agosti, et. al.Donat Agosti ... Alexandros Ioannidis-Pantopikos
09 Oct 2020
Biodiversity Information Science and Standards | VOL. 4

Biodiversity Knowledge Graphs: Time to move up a gear!
Franck Michel ... Julien Kaplan
Biodiversity Information Science and Standards | VOL. 5
Franck Michel, et. al.Franck Michel ... Julien Kaplan
31 Aug 2021
Biodiversity Information Science and Standards | VOL. 5

A framework for integrating biomedical knowledge in Wikidata with open biological and biomedical ontologies and MeSH keywords
Houcemeddine Turki ... Mohamed Ben Aouicha
Heliyon | VOL. 10
Houcemeddine Turki, et. al.Houcemeddine Turki ... Mohamed Ben Aouicha
27 Sep 2024
Heliyon | VOL. 10

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Semantic Search in Legacy Biodiversity Literature: Integrating data from different data infrastructures

Abstract

Talk to us

Similar Papers

More From: Biodiversity Information Science and Standards