Characteristics of scientific Web publications: Preliminary data gathering and analysis

Erik Thorlund Jepsen,Pia Borlund,Piet Seiden,Peter Ingwersen,Lennart Björneborn

doi:10.1002/asi.20079

Abstract

AbstractBecause of the increasing presence of scientific publications on the Web, combined with the existing difficulties in easily verifying and retrieving these publications, research on techniques and methods for retrieval of scientific Web publications is called for. In this article, we report on the initial steps taken toward the construction of a test collection of scientific Web publications within the subject domain of plant biology. The steps reported are those of data gathering and data analysis aiming at identifying characteristics of scientific Web publications. The data used in this article were generated based on specifically selected domain topics that are searched for in three publicly accessible search engines (Google, AllTheWeb, and AltaVista). A sample of the retrieved hits was analyzed with regard to how various publication attributes correlated with the scientific quality of the content and whether this information could be employed to harvest, filter, and rank Web publications. The attributes analyzed were inlinks, outlinks, bibliographic references, file format, language, search engine overlap, structural position (according to site structure), and the occurrence of various types of metadata. As could be expected, the ranked output differs between the three search engines. Apparently, this is caused by differences in ranking algorithms rather than the databases themselves. In fact, because scientific Web content in this subject domain receives few inlinks, both AltaVista and AllTheWeb retrieved a higher degree of accessible scientific content than Google. Because of the search engine cutoffs of accessible URLs, the feasibility of using search engine output for Web content analysis is also discussed.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Characteristics of scientific Web publications: Preliminary data gathering and analysis

Abstract

Talk to us

Similar Papers

More From: Journal of the American Society for Information Science and Technology

Lead the way for us

Journal: Journal of the American Society for Information Science and Technology	Publication Date: Aug 13, 2004
Citations: 31

Similar Papers

Googling for Health Information
Jennifer P D'Auria
Journal of Pediatric Health Care | VOL. 26
Jennifer P D'AuriaJennifer P D'Auria
21 Jun 2012
Journal of Pediatric Health Care | VOL. 26

Beyond TIFF and JPEG2000: PDF/A as an OAIS submission information package container
Yan Han
Library Hi Tech | VOL. 33
Yan HanYan Han
21 Sep 2015
Beyond TIFF and JPEG2000: PDF/A as an OAIS submission information package container
Yan Han

Searching the Web, Continued
Molly Molloy
Science | VOL. 281
Molly MolloyMolly Molloy
10 Jul 1998
Science | VOL. 281

The ProteoRed MIAPE web toolkit: A User-friendly Framework to Connect and Share Proteomics Standards
J Alberto Medina-Aunon ... Miguel A López-García
Molecular & Cellular Proteomics | VOL. 10
J Alberto Medina-Aunon, et. al.J Alberto Medina-Aunon ... Miguel A López-García
19 Jun 2011
Molecular & Cellular Proteomics | VOL. 10

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Characteristics of scientific Web publications: Preliminary data gathering and analysis

Abstract

Talk to us

Similar Papers

More From: Journal of the American Society for Information Science and Technology