Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Data reuse and the open data citation advantage

  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Background. Attribution to the original contributor upon reuse of published data is important both as a reward for data creators and to document the provenance of research findings. Previous studies have found that papers with publicly available datasets receive a higher number of citations than similar studies without available data. However, few previous analyses have had the statistical power to control for the many variables known to predict citation rate, which has led to uncertain estimates of the “citation benefit”. Furthermore, little is known about patterns in data reuse over time and across datasets.Method and Results. Here, we look at citation rates while controlling for many known citation predictors and investigate the variability of data reuse. In a multivariate regression on 10,555 studies that created gene expression microarray data, we found that studies that made data available in a public repository received 9% (95% confidence interval: 5% to 13%) more citations than similar studies for which the data was not made available. Date of publication, journal impact factor, open access status, number of authors, first and last author publication history, corresponding author country, institution citation history, and study topic were included as covariates. The citation benefit varied with date of dataset deposition: a citation benefit was most clear for papers published in 2004 and 2005, at about 30%. Authors published most papers using their own datasets within two years of their first publication on the dataset, whereas data reuse papers published by third-party investigators continued to accumulate for at least six years. To study patterns of data reuse directly, we compiled 9,724 instances of third party data reuse via mention of GEO or ArrayExpress accession numbers in the full text of papers. The level of third-party data use was high: for 100 datasets deposited in year 0, we estimated that 40 papers in PubMed reused a dataset by year 2, 100 by year 4, and more than 150 data reuse papers had been published by year 5. Data reuse was distributed across a broad base of datasets: a very conservative estimate found that 20% of the datasets deposited between 2003 and 2007 had been reused at least once by third parties.Conclusion. After accounting for other factors affecting citation rate, we find a robust citation benefit from open data, although a smaller one than previously reported. We conclude there is a direct effect of third-party data reuse that persists for years beyond the time when researchers have published most of the papers reusing their own data. Other factors that may also contribute to the citation benefit are considered. We further conclude that, at least for gene expression microarray data, a substantial fraction of archived datasets are reused, and that the intensity of dataset reuse has been steadily increasing since 2003.

Similar Papers
  • Research Article
  • Cite Count Icon 19
  • 10.1148/radiol.2016151384
Is There an Association between STARD Statement Adherence and Citation Rate?
  • Feb 2, 2016
  • Radiology
  • Marc Dilauro + 8 more

Purpose To determine if adherence to the Standards for Reporting of Diagnostic Accuracy (STARD) is associated with postpublication citation rates. Materials and Methods A comprehensive search of PubMed, EMBASE, and Cochrane Library databases was performed to identify published articles that have evaluated adherence of diagnostic accuracy studies to the STARD statement. These were included if the number of STARD items reported ("STARD result") could be obtained for each evaluated study. The date of publication, journal impact factor, and citation rate (citations per day) were extracted for the diagnostic accuracy studies. Univariate correlations were performed to identify any association between STARD result, impact factor, and citation rate. Multivariate regression analysis was performed to explore the effect of impact factor on postpublication citation rates. Results The authors were able to obtain the STARD results for 1002 "original" diagnostic accuracy studies from eight different "STARD evaluation" articles. The median impact factor was 3.97 (interquartile range [IQR]: 2.32-6.21), the median STARD result was 15 of 25 items (IQR: 12-18), and the median citation rate was 0.007 citations per day (IQR: 0.0032-0.017). The authors identified a weak positive correlation between STARD result and citation rate (r = 0.096; 95% confidence interval [CI]: 0.034, 0.157), a moderate positive correlation between impact factor and citation rate (r = 0.58; 95% CI: 0.535, 0.617), and a weak positive correlation between impact factor and STARD result (r = 0.13; 95% CI: 0.064, 0.186). Multivariate analysis accounting for journal clustering effects revealed that, when impact factor is partialed out, the positive correlation between citation rate and STARD result does not persist (r = 0.029; 95% CI: -0.033, 0.091). Conclusion There is a positive correlation between completeness of reporting, as evaluated with STARD, and citation rate as well as impact factor. When adjusted for impact factor, the positive correlation between completeness of reporting and citation rate does not persist. (©) RSNA, 2016 Online supplemental material is available for this article.

  • Research Article
  • Cite Count Icon 16
  • 10.1097/bsd.0000000000001303
How Does Open Access Publication Impact Readership and Citation Rates of Lumbar Spine Literature?
  • Mar 3, 2022
  • Clinical Spine Surgery
  • Conor P Lynch + 7 more

This was a retrospective review. The objective of this study was to assess the impact of open access (OA) publication on citation rates and attention scores of literature related to lumbar spine surgery. OA literature allows readers to view full-text manuscripts of research publications free of charge, however, OA publication is often associated with substantial fees for authors. The Altmetric database was searched for articles related to lumbar spine surgery. Title, journal, publication date, Dimensions Citations, Mendeley Readers, Altmetric Attention Score (AAS), number of public mentions, and OA status were collected for each included article. The influence of OA status on Dimensions Citations, Mendeley Readers, and each individual component of the AAS was assessed. To control for journal influence, impact of OA on Dimensions Citations and AAS was separately assessed for each of the top 10 journals contributing the most mentioned articles. The top 25 most cited articles and top 25 articles by AAS were also characterized. A total of 5245 articles were included, of which 2063 were published with OA and 3182 were not. OA status was a significant, independent predictor of AAS and Mendeley Readers (both P <0.001), but not Dimensions Citations ( P =0.422). OA status significantly predicted mentions in news stories ( P =0.003), Twitter posts ( P <0.001), Facebook posts ( P <0.001), and Wikipedia citations ( P =0.011). Of the top 10 contributing journals, OA status significantly predicted Dimensions Citations for European Spine Journal , Journal of Neurosurgery: Spine , and Neurosurgery ( P ≤0.005) and predicted AAS for Spine , European Spine Journal , The Spine Journal , Journal of Neurosurgery: Spine , and Neurosurgery ( P ≤0.017, all). OA status appeared to significantly impact public attention scores, but not citation rates, although these effects did vary based on the journal in which articles were published. Authors may want to consider OA publication based on their target audience and the goal of their research.

  • Research Article
  • Cite Count Icon 6
  • 10.1177/09610006231154529
Allocation of attention to metadata and retrieval functions: Implications for perceived value and open data discovery and reuse
  • Feb 24, 2023
  • Journal of Librarianship and Information Science
  • Ping Wang + 3 more

Metadata and retrieval functions play a vital role in aiding researchers in the discovery and reuse of open data. However, the diversity of metadata elements and retrieval functions poses a challenge to data searchers’ limited attentional resources. This study aims to examine the allocation of attention to metadata elements and retrieval functions and its implications for perceived value and intentions to discover and reuse open data by drawing upon the attentional drift-diffusion model, flow theory, and perceived value literature. An experiment with 48 participants was conducted to explore the proposed relationships. Multiple linear regression analysis was performed to analyze the data. The results suggest that researchers’ attention to high-value functions amplifies the perceived value and motivates data discovery intention. Attention to high-value metadata elements motivates data discovery and reuse intention. In contrast, attention to low-value metadata elements hampers the perceived value and inhibits data discovery and reuse intention. These findings put forward a new lens for exploring the attention mechanisms underlying perceived value, data discovery and reuse intention and highlight the important role of the value of metadata and retrieval functions in attention mechanisms. Additionally, this paper identifies the positive effect of perceived ease of use on users’ intentions to find, evaluate, and access open data. Perceived usefulness positively affects users’ intentions to evaluate open data. However, in contrast to perceived intentions to reuse open data assessed by self-reported measures, perceived value is not a salient motivator of open data reuse intention measured by behavioral indicators. These findings reveal the distinct effects of perceived value on perceived intention and intentional action in data reuse. With these insights, this study develops practical strategies to optimize the design of metadata and retrieval functions in data retrieval systems.

  • Research Article
  • Cite Count Icon 34
  • 10.2519/jospt.2021.10598
The Altmetric Score Has a Stronger Relationship With Article Citations Than Journal Impact Factor and Open Access Status: A Cross-sectional Analysis of 4022 Sport Sciences Articles.
  • Jul 1, 2021
  • Journal of Orthopaedic &amp; Sports Physical Therapy
  • Danilo De Oliveira Silva + 4 more

To assess the relationship of individual article citations in the sport sciences field with (1) Journal Impact Factor, (2) each article's open access status, and (3) Altmetric score components. Cross-sectional. We searched the Web of Science Journal Citation Reports database in the sport sciences category for the 20 journals with the highest 2-year Journal Impact Factor in 2018. We extracted the impact factor for each journal and each article's open access status (yes or no). Between September 2019 and February 2020, we obtained individual citations, Altmetric scores, and details of Altmetric components (eg, number of tweets, Facebook posts, etc) for each article published in 2017. Linear and multiple regression models were used to assess the relationship between the dependent variable (citation number) and the independent variables (article Altmetric score and open access status and Journal Impact Factor). Of the 4022 articles included, the total Altmetric score, Journal Impact Factor, and open access status respectively explained 32%, 14%, and 1% of the variance in article citations (when combined, the variables explained 40% of the variance in article citations). The number of tweets related to an article was the Altmetric component that explained the highest proportion of article citations (37%). Altmetric scores in sport sciences journals have a stronger relationship with number of citations than Journal Impact Factor and open access status do. Twitter may be the best social media platform for promoting a research article. J Orthop Sports Phys Ther 2021;51(11):536-541. Epub 1 Jul 2021. doi:10.2519/jospt.2021.10598.

  • Research Article
  • Cite Count Icon 5
  • 10.13052/jwe1540-9589.197810
Model-based Generation of Web Application Programming Interfaces to Access Open Data
  • Dec 24, 2020
  • Journal of Web Engineering
  • Cesar González-Mora + 3 more

In order to facilitate the reusing of open data from open data platforms’ catalogs, Web Application Programming Interfaces (APIs) are an important mechanism for reusers. However, there is a lack of suitable Web APIs to access data from open data platforms. Moreover, in most cases, the currently available APIs only allow to access catalog’s metadata or to download entire data resources (i.e. coarse-grain access to data), hampering the reuse of data. Therefore, we propose a model-based approach to automatically generate Web APIs from open data. Our generated Web APIs facilitate the access and reuse of specific data (i.e., providing fine-grain or query-level access to data), which will result in many societal and economic benefits such as transparency and innovation. With this approach we address open data publishers which will be able to include a Web API within their data, but also open data reusers in case of missing APIs. This APIfication process, which means the creation of APIs for every available dataset, is based on automatic, generic and standardised generation mechanisms. The performance and functioning of this approach is validated with different datasets, which successfully generates Web APIs that facilitate the reuse of data.

  • Research Article
  • 10.1186/s10195-026-00911-z
Authorship, titles and open access as drivers of citation performance in orthopaedics: a scientometric analysis.
  • Mar 20, 2026
  • Journal of orthopaedics and traumatology : official journal of the Italian Society of Orthopaedics and Traumatology
  • Filippo Migliorini + 7 more

Bibliometric analyses are increasingly used to explore how scientific knowledge is created, disseminated, and perceived. In orthopaedics, research output has expanded rapidly over the past decade, yet the factors determining whether an article achieves wide visibility and scholarly impact remain poorly understood. Beyond the inherent quality of a study, elements such as authorship patterns, title construction, and open access (OA) availability may play an essential role in shaping citation performance. However, evidence in this field is still limited and sometimes contradictory, highlighting the need for large-scale, field-specific analyses. Orthopaedic publications from 2010 to 2020 were identified in Scopus using the keyword 'orthopaedic'. After duplicate removal, 97,806 unique articles were included with complete data on authorship, titles, citation counts, study design, and OA status. Citation rates were normalised per year since publication. Associations between bibliographic features and citation performance were assessed using multiple linear regression, while differences across title styles and study designs were evaluated with comparative statistical testing. Exploratory modelling was performed to identify combinations of authorship and title characteristics linked to the highest predicted citation rates. Larger author teams were associated with higher citation rates (β = 0.108 citations/year per additional author, 95% confidence interval [CI] 0.103-0.114, p < 0.001). OA articles achieved a mean increase of 0.175 citations/year compared with non-OA (p = 0.001). Title length in characters correlated positively with citation rate (β = 0.023 per character, p < 0.001), whereas title length in words showed a negative association (β = -0.183 per word, p < 0.001). The presence of a colon (+0.314 citations/year, p < 0.001) or dash (+0.187, p = 0.001) increased citation performance, while question marks (-0.476, p < 0.001) and all-capital titles (mean 0.71 citations/year) reduced it. Regarding study design, network meta-analyses achieved the highest citation rate (mean 6.64 citations/year), followed by systematic reviews (5.66), meta-analyses (5.08) and narrative reviews (4.81). Randomised controlled trials (3.90) and clinical trials (3.86) performed at an intermediate level, whereas observational studies (2.40), case series (1.79), technical notes (1.33), case reports (0.77), editorials (0.51) and commentaries (0.25) showed consistently lower citation performance (p < 0.0001). In orthopaedic research, collaboration, OA availability and concise, well-structured titles with selected punctuation contribute to higher citation performance, while unconventional title formatting reduces visibility. Although useful for optimising dissemination, ethical authorship practices and rigorous scientific standards remain more critical than citation metrics.

  • Dissertation
  • Cite Count Icon 1
  • 10.4995/thesis/10251/153164
A causal model to explain data reuse in science: a study in health disciplines
  • Sep 18, 2020
  • María Inmaculada Aleixos Borrás

[EN] Investments in data infrastructures, data management, data repositories, and Open Data sharing policies and recommendations are viewed as increasingly important for scientific knowledge production. One of the underlying assumptions justifying these investments is that the more available Open Data becomes, then the greater the possibilities for creating new knowledge that can advance both science and human wellbeing. Yet efforts and investments in Open Data and other ways of data sharing only have value if data are actually reused. Recent scholarly efforts have brought forth some of the challenges and facilitators related to the reuse of data, in order to inform current and future policies and investments. However, despite these efforts, we still do not know why and how some researchers are successful in reusing data, despite the challenges they face, and why some researchers abandon the process of reusing data when facing such challenges. This dissertation aims to fill this gap by focusing on a causal explanation of the data reuse process, which it understands as being nested in broader patterns of researchers' motivations, scientific goals and decision-making strategies.&#13;\n&#13;\nThe dissertation is comprised of three main elements. First, it proposes a heuristic model of the scientific actor, the bounded individual horizon (BIH) model, which understands that, on the one hand, researchers' work and careers are structured by their motivation to produce scientific contributions and rewards systems that prioritizes certain types of contributions. On the other hand, researchers' struggles to achieve their objective of creating new findings that accrue recognition and rewards occur within a frame of limited information and resources, conditioned by multiple institutional, social, and other factors. Second, the study proposes a mechanistic causal theoretical explanation that enables us to understand the data reuse process and its effects (outcomes). The data-reuse mechanism as it is called, enables us to understand how the satisficing behavior that characterizes scientific decision-making applies to the specific conditions and processes of data reuse. Third, a set of ten empirical case studies of data reuse in health research were conducted and are reported in the dissertation. These cases are analyzed and interpreted using the complementary theoretical lenses of the bounded individual horizon and the data-reuse mechanism approaches.&#13;\n&#13;\nThe main findings explain that there is an apparent association between the extent and types of efforts required to reuse data, researchers' contextualized motivations, and broader goal-setting and decision-making frames. Access to data is a necessary condition for the reuse of data, yet is not sufficient for the reuse to happen. Characteristics of available data, including the context of their production, the extent of the preparation and stewarding of these data and their potential value in relation to researchers' motivations to make new scientific claims or generate background knowledge are found to be essential elements for understanding why some data reuse processes persist and succeed, while others do not. The thesis concludes that efforts and investments designed to reap the benefits of data reuse should also be expanded to include training researchers in data reuse, including to efficiently recognize opportunities, navigate the challenges of the reuse process, and be aware of and acknowledge the limitations of the use of secondary data. Without such investments, the promises and expectations linked to emerging data infrastructures, data repositories, data management guidelines and open science practices are argued to be far less likely to reach their full potential.

  • Research Article
  • Cite Count Icon 43
  • 10.1177/03635465221124885
Open Access Articles Garner Increased Social Media Attention and Citation Rates Compared With Subscription Access Research Articles: An Altmetrics-Based Analysis
  • Oct 19, 2022
  • The American Journal of Sports Medicine
  • Amar S Vadhera + 8 more

Background: To better understand the research impact on social media, alternative web-based metrics (Altmetrics) were developed. Open access (OA) publishing, which allows for widespread distribution of scientific content, has become increasingly common in the medical literature. However, the relationship between OA publishing and social media impact remains unclear. Purpose: To compare social media attention and citation rates between OA and subscription access (SA) research articles within the orthopaedic and sports medicine literature. Study Design: Cross-sectional study. Methods: Articles published as either OA or SA in 5 high-impact hybrid orthopaedic journals between January 2019 and December 2019 were analyzed. The primary outcome was the Altmetric Attention Score (AAS), a validated measure of social media attention. Secondary outcomes included citation rates, article characteristics, and the number of shares on social media. Independent t tests and chi-square analyses were used to compare outcomes between OA and SA articles. A multivariable linear regression analysis was performed to determine the association between article type and AAS while controlling for bibliometric characteristics. Results: A total of 2143 articles (246 OA articles, 11.5%; 1897 SA articles, 88.5%) were included. The mean AAS among all OA articles was 62.4 ± 184.6 (range, 0-2032), whereas the mean AAS among all SA articles was 18.4 ± 109.8 (range, 0-3425), representing a statistically significant difference (P < .001). The mean citation rate among OA articles was significantly higher (17.0 ± 22.5; range, 0-139) than that of SA articles (8.6 ± 13.4; range, 0-169) (P < .001). Multivariable linear regression analysis demonstrated that OA status (β = 15.15; P = .044), number of institutions (β = 2.13; P = .023), studies classified as epidemiological investigations (β = 107.40; P < .001), and disclosure of a conflict of interest (β = −11.18; P = .032) were significantly associated with a higher AAS. Conclusion: OA articles resulted in significantly greater AAS and citations in comparison with SA articles. Articles published through the OA option in hybrid journals as well as those with a higher number of institutions, those that disclosed a conflict of interest, and those classified as epidemiological investigations were positively associated with greater AAS in addition to a greater number of citations. The potential for more extensive research dissemination inherent in the OA option may therefore translate into greater reach and social media attention.

  • Research Article
  • Cite Count Icon 1
  • 10.5852/ejt.2025.1004.2971
Open data in publications – non-copyrightability and attribution as drivers for equity, science and innovation
  • Jul 21, 2025
  • European Journal of Taxonomy
  • Jutta Buschbom + 7 more

The mobilization of the wealth of existing biodiversity data is fundamental for the development of large, dynamic and multifaceted datasets and their historical baselines. Biodiversity data need to be reliably and persistently linked to their sources in literary works and in databases. For this purpose, biodiversity data in scholarly publications and their associated repositories have to be transformed to provide findable, accessible, interoperable and reusable resources as a core foundation of interlinked, federated information. Unclear rights and obligations form a substantial obstacle to the reuse and effective interlinking of data, and thus to scientific workflows. Therefore, it is key to understand and implement the legal, ethical and social foundations that are crucial for arriving at informed decisions on the access to and the transformation and reuse of such data.The aim of this article is to provide legal clarity to providers and users of such data, and to give recommendations arising from the analysis of the legal background and of community norms requiring attribution, transparency, and accountability. The goal of the resulting recommendations is to empower the biodiversity sciences and data community, including publishers, authors and users, to apply appropriate legal tools as well as language that will provide legal certainty, thereby accelerating the access to and annotation, extraction and reuse of data contained within publications, both legacy and prospective.The paper is the outcome of a workshop organized during the annual meeting of TDWG, the Biodiversity Information Standards organization, held in Sofia, Bulgaria in October 2022. Its focus was on legal and contractual obligations governing data within works and databases. It also addressed community norms, focusing on attribution as well as ethical and sociocultural principles.

  • Research Article
  • Cite Count Icon 82
  • 10.1108/tg-04-2014-0013
The story of the sixth myth of open data and open government
  • Jan 1, 2015
  • Transforming Government: People, Process and Policy
  • Ann-Sofie Hellberg + 1 more

Purpose – The aim of this paper is to describe a local government effort to realise an open government agenda. This is done using a storytelling approach. Design/methodology/approach – The empirical data are based on a case study. The authors participated in, as well as followed, the process of realising an open government agenda on a local level, where citizens were invited to use open public data as the basis for developing apps and external Web solutions. Based on an interpretative tradition, they chose storytelling as a way to scrutinise the competition process. In this paper, they present a story about the competition process using the story elements put forward by Kendall and Kendall (2012). Findings – The research builds on existing research by proposing the myth that the “public” wants to make use of open data. The authors provide empirical insights into the challenge of gaining benefits from open public data. In particular, they illustrate the difficulties in getting citizens interested in using open public data. Their case shows that people seem to like the idea of open public data, but do not necessarily participate actively in the data reuse process. Research limitations/implications – The results are based on one empirical study. Further research is, therefore, needed. The authors would especially welcome more studies that focus on citizens’ interest and willingness to reuse open public data. Practical implications – This study illustrates the difficulties of promoting the reuse of open public data. Public organisations that want to pursue an open government agenda can use these findings as empirical insights. Originality/value – This paper answers the call for more empirical studies on public open data. Furthermore, it problematises the “myth” of public interest in the reuse of open public data.

  • Research Article
  • Cite Count Icon 1
  • 10.1111/ejss.70103
Facilitating Effective Reuse of Soil Research Data: The BonaRes Repository
  • Mar 1, 2025
  • European Journal of Soil Science
  • Susanne Lachmuth + 5 more

ABSTRACTSoil plays a paramount role in addressing complex challenges related to climate change, the agri‐food system, and ecosystem services. This importance makes soil research data highly relevant for meta‐analysis, research synthesis, modelling, and assessment. As data‐intensive techniques proliferate in studying global change impacts on agricultural systems, effective data management and reuse are essential. Repositories that adhere to the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles are crucial for maximizing the value and efficiency of research data. While publishing in an Open Access repository is necessary for data reusability, it alone is not sufficient. Specialized repositories enhance data reuse potential by addressing discipline‐specific needs through targeted metadata and technical frameworks. The BonaRes Repository was developed for agricultural soil research data and is guided by the FAIR principles, with a focus on data reusability. Here, we introduce the repository's infrastructures and services, including specialized tools for data quality assurance and the management of soil profile as well as long‐term field experiment data. We emphasize the ability of these infrastructures and services to promote data publication and reuse specifically in soil and agricultural sciences. We review examples of data reuse, highlighting their scientific contributions to the understanding of soil and agricultural systems. Finally, we discuss the remaining challenges in achieving FAIR and open soil data publication and reusability. From 2018 to date, the BonaRes Repository has facilitated 815 data publications; 62 papers have reused the published data. Reuse applications range widely—from extracting study site metadata or environmental covariates to reanalysing (meta)data in light of new research questions, to developing scenarios and conducting model calibration and evaluation. A key insight from our review of data reuse is that researchers frequently apply reused data to advance method development. Initiatives such as reciprocal metadata harvesting and integration into larger national and international research data infrastructure will further expand the scope and reuse of the repository's data, including in broader agrosystems science.

  • PDF Download Icon
  • Research Article
  • 10.59490/abe.2016.21.1525
From access to re-use
  • Jan 1, 2016
  • Architecture and the Built Environment
  • Frederika Welle Donker

If data are the building blocks to generate information needed to acquire knowledge and understanding, then geodata, i.e. data with a geographic component (geodata), are the building blocks for information vital for decision-making at all levels of government, for companies and for citizens. Governments collect geodata and create, develop and use geo-information - also referred to as spatial information - to carry out public tasks as almost all decision-making involves a geographic component, such as a location or demographic information. Geo-information is often considered “special” for technical, economic reasons and legal reasons. Geoinformation is considered special for technical reasons because geo-information is multi-dimensional, voluminous and often dynamic, and can be represented at multiple scales. Because of this complexity, geodata require specialised hardware, software, analysis tools and skills to collect, to process into information and to use geoinformation for analyses. Geo-information is considered special for economic reasons because of the economic aspects, which sets it apart from other products. The fixed production costs to create geo-information are high, especially for large-scale geo-information, such as topographic data, whereas the variable costs of reproduction are low which do not increase with the number of copies produced. In addition, there are substantial sunk costs, which cannot be recovered from the market. As such, geo-information shows characteristics of a public good, i.e. a good that is non-rivalrous and non-excludable. However, to protect the high investments costs, re-use of geo-information may be limited by legal and/or technological means such as intellectual property rights and digital rights management. Thus, by making geo-information excludable, it becomes a club good, i.e. a non-rivalrous but excludable good. By claiming intellectual property rights, such as copyright and/or database rights, and restricting (re-)use through licences and licence fees, geo-information can be commercially exploited and used to recover some of the investment costs. Geo-information is considered special for a number of legal reasons. First, as geo-information has a geographic component, e.g. a reference to a location, geoinformation may contain personal data, sensitive company data, environmentally sensitive data, or data that may pose a threat to the national security. Therefore, the dataset may have to be adapted, aggregated or anonymised before it can be made public. Secondly, geo-information may be subject to intellectual property rights. There may be a copyright on cartographic images or database rights on digital information. Such intellectual property rights may be claimed by third parties involved in the information chain, e.g. a private company supplying aerial photography to the National Mapping Authority. The data holder may also claim intellectual property rights to commercially exploit the dataset and recoup some of the vast investment costs made to produce the dataset. Lastly, there may be other (international) legislation or agreements that may either impede or promote publishing public sector information, whereby in some cases, these policies may contradict each other. It has been recognised that to deal with national, regional and global challenges, it is essential that geo-information collected by one level of government or government organisation be shared between all levels of government via a so-called Spatial Data Infrastructure (SDI). The main principles governing SDIs are that data are collected once and (re-)used many times; that data should be easy to discover, access and use; and that data are harmonised so that it is possible to combine spatial data from different sources seamlessly. In line with the SDI governing principles, this dissertation considers accessibility of information to include all these aspects. Accessibility concerns not only access to data, i.e. to be able to view the data without being able to alter the contents but also re-use of data, i.e. to be able to download and/or invoke the data and to share data, including to be able to provide feedback and/or to provide input for co-generated information. Accessibility to public sector geo-information is not only essential for effective and efficient government policy-making but is also associated with realising other ambitions. Examples of these ambitions are a more transparent and accountable government, more citizens’ participation in democratic processes, (co-)generation of solutions to societal problems, and to increase economic value due to companies creating innovative products and services with public sector information as a resource. Especially the latter ambition has been the subject of many international publications stressing the enormous potential economic value of re-use of public sector (geo-) information by companies. Previous research indicated that re-users of public sector information in Europe encountered barriers related to technical, organisational, legal and financial aspects, which was deemed to be the main reason why in Europe the number of value added products and services based on public service information were lagging compared to the United States. Especially the latter two barriers (restrictive licence conditions and high licence fees) were often cited to be the main barriers for reusers in Europe. However, in spite of considerable resources invested by governments to establish spatial data infrastructures, to facilitate data portals and to release public sector information as open data, i.e. without legal and financial restrictions, the expected surge of value added products based on public sector information has not quite eventuated to date and the expected benefits still appear to lag expectations. When this research started a decade ago, the debate around accessibility of public sector information focussed on access policies. Access policies ranged from open access (data available with a minimum of legal restrictions and for no more than marginal dissemination costs) to full cost recovery, whereby all costs incurred in collection, creation, processing, maintenance and dissemination costs to be recovered from the re-users. Most of the public sector bodies in the European Union adhered to a cost recovery policy for allowing re-use of public sector information. In 2003, the European Commission adopted two directives to ensure better accessibility of public sector information Directive 2003/4/EC of the European Parliament and of the Council of 28 January 2003 on public access to environmental information and repealing Council Directive 90/313/EEC, the so-called Access Directive, provided citizens the right of access to environmental information. Citizens should be able to access documents related to the environment via a register, preferably in an electronic form and if a copy of a document was requested, the charges must not exceed marginal dissemination costs. Directive 2003/98/EC of the European Parliament and of the Council of 17 November 2003 on the re-use of public sector information, the co-called PSI Directive, intended to create conditions for a level playing field for all re-users of public sector information. However, the PSI Directive of 2003 left room for public sector organisations to maintain a cost recovery regime with restrictive licence conditions. In spite of these directives, access policies for geographic data were slow to change in most European nations. At the end of the last decade, accessibility of public sector information received two major impulses. The first major impulse was the implementation of Directive 2007/2/ EC of the European Parliament and of the Council of 14 March 2007 establishing an Infrastructure for Spatial Information in the European Community (INSPIRE), the cocalled INSPIRE Directive, established a framework of standardisation rules for the data and publishing via web services, which significantly contributed to the accessibility of public sector geo-information. The second major impulse was the development of open data policies following the Digital Agenda for Europe adopted in 2010 and the USA Open Government Directive of 2009 and the Digital Agenda for Europe of 2010. These two impulses were the main drivers in Europe to start a careful move from cost recovery policies to open access or open data policies and for more public sector information to be made available as open data. Thus, of the four barriers to re-use of public sector information data cited in Chapter 1 (legal, financial, technical and organisational barriers), two barriers should have been lifted to a large degree due to open data. This shift to open data provided an excellent opportunity to test the hypothesis that the main barriers for re-users of public sector information were indeed restrictive licences and high fees as suggested by earlier research. Chapter 2 showed that by 2008, most European Union Member States had transposed and implemented the 2003/98/EC PSI Directive, however, in various ways and with considerable delay. By 2008, the effects of the PSI Directive were only slowly starting to emerge. A number of Member States reviewed their access policies and more public sector information became available for re-use. Some Member States made the information available free-of-charge or reduced their fees significantly. In many cases, where re-use fees were reduced the number of regular re-users increased significantly and total revenue even increased in spite of lower fees. Although the 2007/2/EC INSPIRE Directive paved the way for technical interoperability by providing guidelines for web services and catalogues, neither the INSPIRE Directive nor the PSI Directive had tackled the issue of legal interoperability. Chapter 2 also demonstrated that a major barrier to creating a level playing field for the private sector was the fact that some public sector bodies acted as value added resellers by developing and selling products and services based on their own data. Thus, the level playing field envisioned by the European Commission had not been realised. Chapter 3 researched the aspect of harmonised licences as a first step towards legal interoperability. Earlier research had indicated that one of the biggest barriers for re-users were complex, intransparent and inconsistent licence conditions, especially for re-users wanting to combine data from multiple sources. A survey of licences used by public sector data providers in the Netherlands demonstrated that although there were differences in length and language, there were also many similarities. The conclusion was that the introduction of a licence suite inspired by the Creative Commons concept would be a step towards increased transparency and consistency of geo-information license agreements. This chapter introduced a conceptual model for such a geo-information licence suite, the so-called Geo Shared licences. Both Creative Commons and Geo Shared licence suites enable harmonisation of licence conditions and promote transparency and legal interoperability, especially when re-users combine data from different sources. The Geo Shared licence suite became a serious option for inclusion into the draft version of the INSPIRE Directive as an annex. Unfortunately, the concept of one licence suite for the entire European Union came too early in 2006. The Geo Shared licences were further developed and implemented into the Dutch National Geo Register. In 2009, the European Commission recognised that PSI was the single largest source of information in Europe and the potential for re-use of PSI needed to be highlighted in the digital age. As part of a review of the 2003/98/EC PSI Directive, the European Commission carried out a round of consultations with stakeholders to seek their views on specific issues to be addressed in the future in 2010. In addition, the Commission commissioned a number of studies. These studies included a review of studies on public sector information re-use and related market studies, an assessment of the different models of supply and charging for public sector information and a study on public sector re-user in the cultural sector. The first study, carried out by Graham Vickery in 2011, showed that the overall economic gain from opening up public sector information as a resource for new products and services could be in the order of €40 billion per annum in the European Union. Both the Vickery Report and the second study, the so-called POPSIS Study, showed that for most public sector data providers their revenues from licence fees were relatively low in comparison to their total budget. After the evaluation, Directive 2013/37/EU of the European Parliament and of the Council of 26 June 2013 amending Directive 2003/98/EC on the re-use of public sector information was adopted and came into force on 17 July 2013. Chapter 4 described the main changes of the 2013/37/EU Amended PSI Directive, including the recommendation to employ open data licences. This chapter continued with a review of the various open data licences in use in Europe and analysed their interoperability. Although adoption of open data licences for public sector information should have addressed legal interoperability barriers for re-users, in practice, the different types of open data licences might not be so interoperable after all. Effectively, only a public domain declaration, such as a Creative Commons Zero (CC0) declaration, is suitable for open data re-users requiring with cross-border data sets and that such a public domain declaration is published in a prominent place to remove uncertainty for re-users. Without a public domain declaration, re-use of open data is still impeded as re-users are loathe to invest time into the development of value added products or services when it is uncertain if and which restrictions may be applicable and what the impact may be on their product or service. This dissertation also researched the financial and economic aspects of public sector information accessibility. Chapters 1 and 2 indicated that a cost recovery regime for dissemination of public sector information provided a financial barrier for private sector re-users because the fees charged were perceived to be too high. However, in 2008, there were still many advocates for maintaining a cost recovery regime. Especially public sector bodies that are not funded by the national Treasury, the socalled self-funding agencies, needed revenue from data sales to cover a substantial part of their operational costs. A sustainable source of revenue was viewed as essential to maintain the data at an adequate level, and to ensure actuality and continuity. Chapter 5 explored the potential business models and pricing mechanisms for public sector INSPIRE web services. Although, depending on the type of web service, and type of re-user, there might have been an argument for employing a subscription model as a pricing mechanism, business models based on generating revenue from public sector information would not be viable in the long run and were not in the spirit of the INSPIRE Directive. This research concluded that public sector information web services employing different pricing regimes were counterproductive to achieving financial interoperability. In Chapter 6, business models for public sector data providers were revisited, this time from an open data perspective. Government agencies, including self-funding government agencies are under increasing pressure to implement open data policies. This chapter analysed the business models of self-funding agencies either already providing open data or under pressure to provide (some) open data in the near future. The analysis showed which adaptions might be necessary to ensure the long-term availability of high quality open data and the long-term financial sustainability of self-funding agencies. The case studies confirmed that providing (raw) open data does not necessarily lead to losses in revenue in the long term as long as the organisation has enough flexibility to adapt its role in the information value chain, especially when revenue from licence fees represents only a relative small part of their total budget. The case studies indicated that switching to open data has resulted in internal efficiency gains. In practice, it is difficult to isolate and quantify the internal efficiency gains that are solely attributable to open data as the researched organisations continuously implement efficiency measures. However, the reported decreases in internal and external transaction costs due to open data are in line with the case study carried out in Chapter 7. Open data also provided an excellent opportunity to assess the effects of open data ex ante as baseline measurements could be carried out. To develop both quantitative and qualitative indicators to assess the success of a policy change is a challenge for open data initiatives. In Chapter 7, a model to assess the effects on the organisation of an open data provider was developed. Liander, a private energy network administrator mandated with a public task, planned to publish some of their datasets as open data in the autumn of 2013. This offered an excellent opportunity to apply the developed assessment model to provide an insight into internal, external, and relational effects on Liander. A benchmark was carried out prior to release of open data and a follow-up measurement one year later. The benchmark provided an insight into the then work processes and into the preparations required to implement open data. The follow-up monitor indicated that Liander open data are used by a wide range of users and have had a positive effect on the development of apps to aid energy savings. However, it remains a challenge to quantify the societal effects of such apps. The follow-up monitor also indicated that regular re-users of Liander data used the open data to improve existing applications and work processes rather than to create new products. The case study demonstrated that private energy companies could successfully release open data. The case study also showed that Liander served as a best-practice case for open data and had a flywheel effect on companies within the same sector. By 2015, nearly all energy network administrators had published similar open data. The monitoring model developed in this project was assessed to be suitable to monitor the open data effects on the organisation of the data provider. The assessment model developed and tested in Chapter 7 proved to be suitable to monitor the effects of open data on organisational level. However, to provide a more complete picture of the effects of open data and to assess if there are other barriers for re-users, a more holistic approach was required to assess the maturity of open data. Therefore, a holistic open data assessment framework addressing the supplier side, the governance side, and the user side of the open data was developed and applied to the Dutch open data infrastructure in Chapter 8. This Holistic Open Data Maturity Assessment Framework was used to evaluate the State of the Open Data Nation in the Netherlands and to provide valuable information on (potential) bottlenecks. The framework showed that geographic data scored significantly better than other types of government data. The standardisation and implementation rules laid down by INSPIRE Directive framework appear to have been a catalyst for moving geographic data to a higher level of maturity. The maturity assessment framework provided Dutch policy makers with useful inputs for further development of the open data ecosystem and development of well-founded strategies that will ensure the full potential of open data will be reached. Since the publication of the State of the Open Data Nation in 2014, a number of the recommendations have already been implemented. This dissertation demonstrated that many aspects that should facilitate accessibility, such as standardised metadata, have already been addressed for geodata. This research also showed that for other types of data, there is still a long way to go. There is a growing demand for other types of data, such as financial data and healthcare data. Public sector organisations holding such types of data need hands-on guidelines to enable publication of their datasets, preferably as open data. However, data published as open data are forever and cannot be recalled. Therefore, the decision to publish public sector data as open data is complex: datasets are often of a heterogeneous nature and may contain microdata (data that quantify observations or facts, such as data collected during surveys) Although microdata may not necessarily contain personal data, the datasets will probably have to be processed before publication to address confidentiality and data quality issues. In addition, there is a tension between open data and protection of personal data. The big question remains to which level the data need to be aggregated and/or anonymised to ensure protection of personal data now and in the future, and at the same time keeping sufficient significance to be re-usable. Another issue that needs further research is data-ownership of sensor data and co-created data. Increasingly, sensor data generated by e.g. smart phones, smart energy meters and traffic sensors are collected by the public sector and the private sector and become part of a big data ecosystem. In addition, public sector organisations cooperate with other public sector organisations and the private sector to create information from their data, so-called co-created information. Citizens also collect data or complement information on a voluntary basis, e.g. bird counts data. Co-created information will become more commonplace in the coming decades, as will the contribution of sensor data to a big data ecosystem. However, the aspect of who owns the data in which part of the information value chain has not been researched. Uncertainty related to third party rights will pose a barrier to publishing open data. Therefore, the aspect of data-ownership for sensor data and for co-created data should be further researched.

  • Research Article
  • Cite Count Icon 99
  • 10.1177/0363546520903703
What Is the Predictive Ability and Academic Impact of the Altmetrics Score and Social Media Attention?
  • Feb 28, 2020
  • The American Journal of Sports Medicine
  • Kyle N Kunze + 6 more

Background: Citation rate and journal impact factor have traditionally been used to assess research impact; however, these may fail to represent impact beyond the sphere of academics. Given that social media is now used to disseminate research, alternative web-based metrics (altmetrics) were recently developed to better understand research impact on social media. However, the relationship between altmetrics and traditional bibliometrics in orthopaedic literature is poorly understood. Purpose: To (1) assess the extent that altmetrics correlate with traditional bibliometrics and (2) identify publication characteristics that predict greater altmetrics scores. Study Design: Cross-sectional study. Methods: Articles published in The American Journal of Sports Medicine (AJSM), The Journal of Bone and Joint Surgery, Clinical Orthopaedics and Related Research, Acta Orthopaedica, and Knee Surgery, Sports Traumatology, Arthroscopy between January 2016 and December 2016 were analyzed. Among the extracted publication characteristics were journal, number of authors, geographic region of origin, highest degree of first author, study subject and design, sample size, conflicts of interest, and level of evidence; number of references, institutions, citations, tweets, Facebook mentions, and news mentions; and Altmetric Attention Score (AAS). Multivariate regressions were used to determine (1) publication characteristics predictive of AAS and social media attention (mentions on Twitter, Facebook, and the news) and (2) the relationship between AAS and citation rate. Results: A total of 496 published articles were included, with a mean AAS of 8.6 (SD, 31.7; range, 0-501) and a mean citation rate of 15.0 (SD, 16.1; range, 0-178). Articles in AJSM (β = 19.9; P < .001), publications from North America (β = 8.5; P = .033), and studies concerning measure validation/reliability (β = 25.5; P = .004) were independently associated with higher AAS. Greater AAS score significantly predicted a greater citation rate (β = 0.16; P < .0001). The citation rate was an independent predictor of greater social media attention on Twitter, Facebook, and the news (odds ratio range, 1.02-1.03; P < .05 all). Conclusion: AAS had a significant positive association with citation rates of articles in 5 high-impact orthopaedic journals. Articles in AJSM, studies concerning measure validation and reliability, and publications from North America were positively associated with greater AAS. A greater number of citations was consistently associated with publication attention received on social media platforms.

  • Research Article
  • Cite Count Icon 24
  • 10.1093/asj/sjz336
Citation Skew in Plastic Surgery Journals: Does the Journal Impact Factor Predict Individual Article Citation Rate?
  • Nov 20, 2019
  • Aesthetic Surgery Journal
  • Malke Asaad + 5 more

Citation skew refers to the unequal distribution of citations to articles published in a particular journal. We aimed to assess whether citation skew exists within plastic surgery journals and to determine whether the journal impact factor (JIF) is an accurate indicator of the citation rates of individual articles. We used Journal Citation Reports to identify all journals within the field of plastic and reconstructive surgery. The number of citations in 2018 for all individual articles published in 2016 and 2017 was abstracted. Thirty-three plastic surgery journals were identified, publishing 9823 articles. The citation distribution showed right skew, with the majority of articles having either 0 or 1 citation (40% and 25%, respectively). A total of 3374 (34%) articles achieved citation rates similar to or higher than their journal's IF, whereas 66% of articles failed to achieve a citation rate equal to the JIF. Review articles achieved higher citation rates (median, 2) than original articles (median, 1) (P < 0.0001). Overall, 50% of articles contributed to 93.7% of citations and 12.6% of articles contributed to 50% of citations. A weak positive correlation was found between the number of citations and the JIF (r = 0.327, P < 0.0001). Citation skew exists within plastic surgery journals as in other fields of biomedical science. Most articles did not achieve citation rates equal to the JIF with a small percentage of articles having a disproportionate influence on citations and the JIF. Therefore, the JIF should not be used to assess the quality and impact of individual scientific work.

  • Peer Review Report
  • 10.7554/elife.82498.sa1
Decision letter: Tracing the path of 37,050 studies into practice across 18 specialties of the 2.4 million published between 2011 and 2020
  • Oct 11, 2022
  • Alan Schroeder + 1 more

Comprehensive decade-long analysis of point-of-care resources reveals insights into how clinical research makes its way into practice across different medical specialties and through time to enable us to identify potential translational bottlenecks in the pathway from research to practice.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant