MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format

Karin Verspoor,Juilee Thakar,Thomas Dandekar,Florencio Pazos,Zeeshan Ahmed,M Julius Hossain

doi:10.5256/f1000research.7898.r11637

Karin Verspoor, Juilee Thakar + Show 4 more

https://doi.org/10.5256/f1000research.7898.r11637

Copy DOI

Journal: F1000Research	Publication Date: Apr 11, 2018
Citations: 2	License type: CC BY 4.0

Affiliation: University of Melbourne, University of Würzburg

Abstract

Published scientific literature contains millions of figures, including information about the results obtained from different scientific experiments e.g. PCR-ELISA data, microarray analysis, gel electrophoresis, mass spectrometry data, DNA/RNA sequencing, diagnostic imaging (CT/MRI and ultrasound scans), and medicinal imaging like electroencephalography (EEG), magnetoencephalography (MEG), echocardiography (ECG), positron-emission tomography (PET) images. The importance of biomedical figures has been widely recognized in scientific and medicine communities, as they play a vital role in providing major original data, experimental and computational results in concise form. One major challenge for implementing a system for scientific literature analysis is extracting and analyzing text and figures from published PDF files by physical and logical document analysis. Here we present a product line architecture based bioinformatics tool ‘Mining Scientific Literature (MSL)’, which supports the extraction of text and images by interpreting all kinds of published PDF files using advanced data mining and image processing techniques. It provides modules for the marginalization of extracted text based on different coordinates and keywords, visualization of extracted figures and extraction of embedded text from all kinds of biological and biomedical figures using applied Optimal Character Recognition (OCR). Moreover, for further analysis and usage, it generates the system’s output in different formats including text, PDF, XML and images files. Hence, MSL is an easy to install and use analysis tool to interpret published scientific literature in PDF format.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format

Abstract

Talk to us

Similar Papers

More From: F1000Research

Lead the way for us

Similar Papers

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format
Zeeshan Ahmed ... Thomas Dandekar
F1000Research | VOL. 4
Zeeshan Ahmed, et. al.Zeeshan Ahmed ... Thomas Dandekar
12 Apr 2017
F1000Research | VOL. 4

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format.
Zeeshan Ahmed ... Thomas Dandekar
F1000Research | VOL. 4
Zeeshan Ahmed, et. al.Zeeshan Ahmed ... Thomas Dandekar
04 Apr 2018
F1000Research | VOL. 4

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format.
Zeeshan Ahmed ... Thomas Dandekar
F1000Research | VOL. 4
Zeeshan Ahmed, et. al.Zeeshan Ahmed ... Thomas Dandekar
16 Dec 2015
F1000Research | VOL. 4

MSL: Mining published scientific literature for the extraction and classification of text and images to support IR capabilities
Ahmed Zeeshan ... Zeeshan Saman
Frontiers in Neuroinformatics | VOL. 10
Ahmed Zeeshan, et. al.Ahmed Zeeshan ... Zeeshan Saman
01 Jan 2015
Frontiers in Neuroinformatics | VOL. 10

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MSL: Facilitating automatic and physical analysis of published scientific literature in PDF format

Abstract

Talk to us

Similar Papers

More From: F1000Research