MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format.

Zeeshan Ahmed,Thomas Dandekar

doi:10.12688/f1000research.7329.3

Zeeshan Ahmed, Thomas Dandekar

Open Access

https://doi.org/10.12688/f1000research.7329.3

Copy DOI

Journal: F1000Research	Publication Date: Apr 4, 2018
Citations: 1	License type: CC BY 4.0

Affiliation: Jackson Laboratory, University of Würzburg

Abstract

Published scientific literature contains millions of figures, including information about the results obtained from different scientific experiments e.g. PCR-ELISA data, microarray analysis, gel electrophoresis, mass spectrometry data, DNA/RNA sequencing, diagnostic imaging (CT/MRI and ultrasound scans), and medicinal imaging like electroencephalography (EEG), magnetoencephalography (MEG), echocardiography (ECG), positron-emission tomography (PET) images. The importance of biomedical figures has been widely recognized in scientific and medicine communities, as they play a vital role in providing major original data, experimental and computational results in concise form. One major challenge for implementing a system for scientific literature analysis is extracting and analyzing text and figures from published PDF files by physical and logical document analysis. Here we present a product line architecture based bioinformatics tool 'Mining Scientific Literature (MSL)', which supports the extraction of text and images by interpreting all kinds of published PDF files using advanced data mining and image processing techniques. It provides modules for the marginalization of extracted text based on different coordinates and keywords, visualization of extracted figures and extraction of embedded text from all kinds of biological and biomedical figures using applied Optimal Character Recognition (OCR). Moreover, for further analysis and usage, it generates the system's output in different formats including text, PDF, XML and images files. Hence, MSL is an easy to install and use analysis tool to interpret published scientific literature in PDF format.

Highlights

There has been an enormous increase in the amount of the scientific literature in the last decades[1]
We developed Mining Scientific Literature (MSL)’s Text module, which is capable of processing Portable Document Format (PDF) files with single, double or multiple columns
In the case of text extraction we observed that the text was in reading order when using manuscripts from F1000Research and IEEE but text was without spaces in the manuscript from PLOS and with additional lines and extra spaces in the manuscript from Hindawi

Summary

Introduction

There has been an enormous increase in the amount of the scientific literature in the last decades[1]. Most published scientific literature is available in Portable Document Format (PDF), a very common way for exchanging printable documents. This makes it all-important to extract text and figures from the PDF files to implement an efficient Natural Language Processing (NLP) based search application. PDF is only rich in displaying and printing but requires explicit efforts in the extraction of information, which significantly impacts the search and retrieval capabilities[2]. Due to this reason several document analysis based tools have been developed for physical and logical document structure analysis of this file type

Objectives

Methods

Results

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: F1000Research

Lead the way for us

Similar Papers

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format
Zeeshan Ahmed ... Thomas Dandekar
F1000Research | VOL. 4
Zeeshan Ahmed, et. al.Zeeshan Ahmed ... Thomas Dandekar
12 Apr 2017
F1000Research | VOL. 4

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format
Karin Verspoor ... M Julius Hossain
F1000Research | VOL. 4
Karin Verspoor, et. al.Karin Verspoor ... M Julius Hossain
11 Apr 2018
F1000Research | VOL. 4

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format.
Zeeshan Ahmed ... Thomas Dandekar
F1000Research | VOL. 4
Zeeshan Ahmed, et. al.Zeeshan Ahmed ... Thomas Dandekar
16 Dec 2015
F1000Research | VOL. 4

MSL: Mining published scientific literature for the extraction and classification of text and images to support IR capabilities
Ahmed Zeeshan ... Dandekar Thomas
Frontiers in Neuroinformatics | VOL. 10
Ahmed Zeeshan, et. al.Ahmed Zeeshan ... Dandekar Thomas
01 Jan 2015
Frontiers in Neuroinformatics | VOL. 10

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: F1000Research