Metadata Information Research Articles

Classifying scientific publications according to Field-of-Science taxonomies is of crucial importance, powering a wealth of relevant applications including Search Engines, Tools for Scientific Literature, Recommendation Systems, and Science Monitoring. Furthermore, it allows funders, publishers, scholars, companies, and other stakeholders to organize scientific literature more effectively, calculate impact indicators along Science Impact pathways and identify emerging topics that can also facilitate Science, Technology, and Innovation policy-making. As a result, existing classification schemes for scientific publications underpin a large area of research evaluation with several classification schemes currently in use. However, many existing schemes are domain-specific, comprised of few levels of granularity, and require continuous manual work, making it hard to follow the rapidly evolving landscape of science as new research topics emerge. Based on our previous work of scinobo, which incorporates metadata and graph-based publication bibliometric information to assign Field-of-Science fields to scientific publications, we propose a novel hybrid approach by further employing Neural Topic Modeling and Community Detection techniques to dynamically construct a Field-of-Science taxonomy used as the backbone in automatic publication-level Field-of-Science classifiers. Our proposed Field-of-Science taxonomy is based on the OECD fields of research and development (FORD) classification, developed in the framework of the Frascati Manual containing knowledge domains in broad (first level(L1), one-digit) and narrower (second level(L2), two-digit) levels. We create a 3-level hierarchical taxonomy by manually linking Field-of-Science fields of the sciencemetrix Journal classification to the OECD/FORD level-2 fields. To facilitate a more fine-grained analysis, we extend the aforementioned Field-of-Science taxonomy to level-4 and level-5 fields by employing a pipeline of AI techniques. We evaluate the coherence and the coverage of the Field-of-Science fields for the two additional levels based on synthesis scientific publications in two case studies, in the knowledge domains of Energy and Artificial Intelligence. Our results showcase that the proposed automatically generated Field-of-Science taxonomy captures the dynamics of the two research areas encompassing the underlying structure and the emerging scientific developments.

Read full abstract

Despite the importance of calanoid copepods to healthy ecosystem functioning of the Arctic Ocean and Subarctic Seas, many aspects of their biogeography, particularly in winter months, remain unresolved. At the same time, online databases that digitize species distribution records are growing in popularity as a tool to investigate ecological patterns at macro scales. The value of such databases for Calanus research requires investigation - the long history of Calanus sampling holds promise for such databases, while conditions at high latitudes may impose limits through spatial and temporal biases. We collated records of three Calanus species (C. finmarchicus, C. glacialis, and C. hyperboreus) from the Ocean Biodiversity Information System (OBIS) and the Global Biodiversity Information Facility (GBIF) providing over 230,000 unique records spanning 150 years and over 100 individual datasets. After quality control and cleaning, the latitudinal and vertical distribution of occurrences were explored, as well as the completeness of informative metadata fields. Calanus sampling was found to be temporally and spatially biased towards surfacemost layers (&lt;10m) in spring and summer. Only 3.5% of records had an average collection depth ≥400m, approximately half of these in months important for diapause. Just over 40% of records lacked associated information on sampling protocol while 11% of records lacked life-stage information. OBIS data contained fields for maximum and minimum collection depth and so were subset into discrete “shallow summer” and “deep winter” life cycle phases and matched to sea-ice and temperature conditions. 23% of OBIS records north of 66° latitude were located in regions of seasonal sea-ice presence and occurrences show species-specific thermal optima during the shallow summer period. The collection depth of C. finmarchicus was significantly different to C. hyperboreus during the deep winter. Overall, online databases contain a vast number of Calanus records but sampling biases should be acknowledged when they are used to investigate patterns of biogeography. We advocate efforts to integrate additional data sources within online portals. Particular gaps to be filled by existing or future collections are (i) widening the spatial extent of sampling during spring/summer months, (ii) increasing the frequency of sampling during winter, particularly at depths below 400m, and (iii) improving the quality, quantity and consistency of metadata reporting.

Read full abstract

Metadata Information Research Articles

Related Topics

Articles published on Metadata Information

Application of AI-Helped Image Classification of Fish Images: An iDigBio dataset example

Characterizing Errors Using Satellite Metadata for Eco‐Hydrological Model Calibration

Collection insight and interconnectivity through artificial intelligence image analysis: A collaboration with the National Archives of Estonia

Near real-time predictions of renewable electricity production at substation level via domain adaptation zero-shot learning in sequence

Novel machine learning based authentication technique in VANET system for secure data transmission

Tissue proteomics repositories for data reanalysis.

Comparison of Corecorded Analog and Digital Systems for Characterization of Responses and Uncertainties

Book Review: Metadata for Digital Collections, Second Edition

Epoxy: ACID Transactions across Diverse Data Stores

How Do Users Examine Online Messages to Determine If They Are Credible? An Eye-Tracking Study of Digital Literacy, Visual Attention to Metadata, and Success in Misinformation Identification

Enhancing accessibility for the blind and visually impaired: Presenting semantic information in PDF tables

A Qualitative Analysis on Pesantren Economic

MSdb: An integrated expression atlas of human musculoskeletal system

Co-Clinical Imaging Metadata Information (CIMI) for Cancer Research to Promote Open Science, Standardization, and Reproducibility in Preclinical Imaging.

SCINOBO: a novel system classifying scholarly communication in a dynamically constructed hierarchical Field-of-Science taxonomy.

Forensik Video Pada CCTV Menggunakan Framework Generic Computer Forensics Investigation Model (GCFIM)

Approaches and tools for user-driven provenance and data quality information in spatial data infrastructures

Assessing key influences on the distribution and life-history of Arctic and boreal Calanus: are online databases up to the challenge?

Recent flex power changes

Giving Historical Photographs a New Perspective: Introducing Camera Orientation Parameters as New Metadata in a Large-Scale 4D Application

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Metadata Information Research Articles

Related Topics

Articles published on Metadata Information

Application of AI-Helped Image Classification of Fish Images: An iDigBio dataset example

Characterizing Errors Using Satellite Metadata for Eco‐Hydrological Model Calibration

Collection insight and interconnectivity through artificial intelligence image analysis: A collaboration with the National Archives of Estonia

Near real-time predictions of renewable electricity production at substation level via domain adaptation zero-shot learning in sequence

Novel machine learning based authentication technique in VANET system for secure data transmission

Tissue proteomics repositories for data reanalysis.

Comparison of Corecorded Analog and Digital Systems for Characterization of Responses and Uncertainties

Book Review: Metadata for Digital Collections, Second Edition

Epoxy: ACID Transactions across Diverse Data Stores

How Do Users Examine Online Messages to Determine If They Are Credible? An Eye-Tracking Study of Digital Literacy, Visual Attention to Metadata, and Success in Misinformation Identification

Enhancing accessibility for the blind and visually impaired: Presenting semantic information in PDF tables

A Qualitative Analysis on Pesantren Economic

MSdb: An integrated expression atlas of human musculoskeletal system

Co-Clinical Imaging Metadata Information (CIMI) for Cancer Research to Promote Open Science, Standardization, and Reproducibility in Preclinical Imaging.

SCINOBO: a novel system classifying scholarly communication in a dynamically constructed hierarchical Field-of-Science taxonomy.

Forensik Video Pada CCTV Menggunakan Framework Generic Computer Forensics Investigation Model (GCFIM)

Approaches and tools for user-driven provenance and data quality information in spatial data infrastructures

Assessing key influences on the distribution and life-history of Arctic and boreal Calanus: are online databases up to the challenge?

Recent flex power changes

Giving Historical Photographs a New Perspective: Introducing Camera Orientation Parameters as New Metadata in a Large-Scale 4D Application