The health informatics centre: a safe haven and trusted research environment enabling world-leading research
IntroductionThe Health Informatics Centre (HIC) is a regional Scottish Safe Haven dedicated to secure data management, ensuring its integrity, confidentiality, and availability through a robust information governance framework.MethodsAs a data processor, HIC is responsible for the secure curation, storage, and provision of research data extracts. Research-ready data are made available to approved researchers via our customisable cloud Trusted Research Environment (TRE).ResultsThe available granular data spans over 20 years, includes 2.1 million of the Scottish population, and HIC offers more than 170 datasets, with the most commonly used published with a digital object identifier. Data sources include clinical, hospital, laboratory, imaging, and research datasets which can be linked to new, and existing datasets. The data is quality-assured and released as project-specific extracts, ensuring robust privacy protection and research readiness. HIC’s infrastructure is secure-by-design and supports high-performance computing, advanced data analytics, and is customisable to researcher’s needs.ConclusionHIC has a long history of supporting a wide range of data-led research projects as a trusted and capable partner. At the time of publication 175 projects across academia, the NHS, and public sector organisations are active within HIC. Through adaptability, innovation and investment in people and infrastructure we have established a sustainable model which will continue to meet future needs and demands from world-leading research with sensitive data.
- Abstract
- 10.23889/ijpds.v1i1.258
- Apr 18, 2017
- International Journal of Population Data Science
ABSTRACT
 Objectives
 
 Design and implement an architecture for managing unconsented DICOM imaging
 Maintain sufficient data to define research cohorts when data quality is unknown
 Perform project-level linkage and extraction into a Safe Haven (SH) environment
 Extract large image volumes for multiple projects with limited storage constraints
 Provide applications for an imaging research workflow within the SH environment
 Serve as a prototype for the Farr/NHS Scotland project to create a research dataset from Scotland’s national PACS
 
 ApproachThe software architecture builds on the Research Data Management Platform (RDMP) developed at Dundee’s Health Informatics Centre (HIC) within Farr@Dundee. The RDMP provides core services common to loading any dataset, with configuration and extensibility points for dataset-specific implementations. This architecture augments the RDMP with scalable micro-services performing peripheral functions.
 Images are sourced from the local PACS server in Ninewells Hospital and cached securely within HIC using an implementation for the RDMP with a custom server to query/retrieve data.
 Data stored in the catalogue should be anonymous, according to the Scottish SH model. The imaging dataset is poorly understood, with several potentially identifiable free-text fields which may contain information required for defining suitable research cohorts. The load process only permits verified metadata fields into the anonymised catalogue; a Mongo database stores other data for later analysis, should a field subsequently be required for cohort definition.
 A DICOM extraction implementation is provided, using DICOM Confidential for anonymisation and a project-specific remapping of DICOM GUIDs.
 Two provisioning methods have been designed. A basic copy when sufficient storage is available, and a more sophisticated method using a custom filesystem to provide separate project-specific views onto shared image files.
 ResultsA full end-to-end solution has been developed, from initial caching through to provisioning anonymised images. Two imaging cohorts have been loaded, one with over 5000 studies. NHS Tayside CT and MR data since 2008 is currently being loaded.
 Two projects have had anonymised extracts released using the ‘copy’ method. The custom filesystem method has been developed and tested with limited amounts of data.
 This work has highlighted anonymisation, cohort creation and SH issues which require further exploration.
 ConclusionA production system for securely providing linked DICOM imaging to researchers has been implemented, serving as a testbed for a national system which will provide a unique population-level resource for researchers.
- Conference Article
1
- 10.14236/ewic/his2014.1
- Jan 1, 2014
- Electronic workshops in computing
Clinical datasets are the most critical resources or assets in the repository of Electronic Health Records (EHRs) and their quality gains competitive advantages in translational research. Accurate, reliable, and consistent representation of clinical datasets are essential for answering key research questions. However, a major issue with carrying out research on routinely collected primary care datasets is that they are often not fit-for-purpose or research-ready. It often takes months (if not years) for researchers to clean and transform clinical datasets for meaningful translational research. Profiling clinical datasets provides a proactive approach to examining and understanding the content, context and structure of source system data. The objective of this study was to develop a profiling dashboard to monitor, measure, assess, and improve the quality of clinical datasets hosted and maintained by the Health Informatics Centre (HIC) at the University of Dundee. Preliminary results indicated that the dashboard affords the flexibility to perform objective assessment of data quality, in terms of accessibility, accuracy, appropriate amount of data, completeness, and consistency.
- Research Article
- 10.3897/biss.3.38263
- Jul 17, 2019
- Biodiversity Information Science and Standards
Since 2015, the Natural History Museum London has made its research and collections data available through its Data Portal (https://data.nhm.ac.uk). This website provides free and open access to important research datasets as well as digitised objects from the Museum's specimen collection. The Data Portal currently has over 4.2 million records from the specimen collection and a further 5.5 million records from other research datasets. Since 2015, more than 250 scientific publications have cited data from the Data Portal, either directly or through aggregators such as the Global Biodiversity Information Facility (GBIF), although there are many more citations than it is currently possible to track. Users can download data from the Portal and are encouraged to cite the source, however, there is currently no way for users to cite subsets of the data returned through a query, nor a way to persistently identify the data subset they are citing. This is a common issue with scientific data put online, particularly when the cited data changes frequently, such as is the case with the Museum's specimen collection, which grows constantly as more of the collection is digitised. This poster outlines a new approach that has been designed to meet the Research Data Alliance's (RDA) Working Group on Data Citation recommendations on citing evolving data (Rauber et al. 2015). This is achieved by implementing a fully versioned search framework, ensuring that all modifications to records are tracked and the version timestamp of each modification is combined with the data into the search index. When users search and download data from the Portal, Digital Object Identifiers (DOI) are minted for unique searches at exact versions allowing the dynamic, repeated retrieval of data at any version timestamp, without storing the results. Combining the versioning information into the search index also allows queries against historical data. By persistently identifying query results in this fashion, researchers can cite data precisely and have confidence that although the data may change after they use it, users of their work will be able to access the data as it looked when they studied it originally. This should also encourage the systematic use of citations, making it easier to track both the usage and impact of research and collections datasets.
- Research Article
28
- 10.3109/17538157.2014.892491
- Mar 20, 2014
- Informatics for Health and Social Care
Background: Combining routinely collected health and social care data on older people is essential to advance both service delivery and research for this client group. Little data is available on how to combine health and social care data; this article provides an overview of a successful data linkage process and discusses potential barriers to executing such projects.Methods and results: We successfully obtained and linked data on older people within Dundee from three sources: Dundee Social Work Department database (30 000 individuals aged 65 years and over), healthcare data held on NHS Tayside patients by the Health Informatics Centre (400 000 individuals), Dundee, and the Dundee of Medicine for the Elderly rehabilitation database (4300 individuals). Data were linked, anonymized and transferred to a Safe Haven environment to ensuring confidentiality and strict access control. Challenges were faced around workflows, culture and documentation. Exploiting the resultant data set raises further challenges centered on database documentation, understanding the way data were collected, dealing with missing data, data validity and collection at different time periods.Conclusion: Routinely collected health and social care data sets can be linked, but significant process barriers must be overcome to allow successful linkage and integration of data and its full exploitation.
- Dissertation
- 10.53846/goediss-6205
- Jan 1, 2017
Research data occurs in all scientific experiments, computer simulations, observations or as a derivation from other datasets, literature or publications. As a subset of the general concept of digital data, it is classified through its distinct state and its origin. Enriched with descriptive metadata, research data serves as a foundation for discoveries and publishing results in various formats. For citing and linking specific research datasets and publications, unique and persistent identification is necessary. Today, this is realized by Persistent Identifier (PID) systems that provide stable identification for digital entities and an optional annotation by descriptive metadata. Moreover, PID systems abstract the current network location of data in order to anticipate changes in its network location, owed to alternating Uniform Resource Locators (URL) on the World Wide Web (WWW). Applying these concepts, PID systems have tagged billions of research datasets and publications over the past 20 years. On these foundations, the Handle PID system, known from the Digital Object Identifier (DOI) system, provides reliable access to digital publications and research data to the whole scientific community. While the architecture of the Handle system itself, which depends on fixed network locations, was designed with farsightedness, additional end-user services for PID resolution and management have introduced critical weak spots that can be discovered by comprehensively reviewing the current state-of-the-art. This thesis focuses on the adaption of location-independent network paradigms which have shown encouraging results when applied to several problems in the domain of decentralized network infrastructures in PID systems. Our first approaches aim at evolving the Handle system design into a self-adjusting system for all major infrastructure services that does not depend on fixed network locations. We tackle this by incorporating strategies and techniques from location-independent network paradigms originating from the current research branch of Named Data Networking (NDN). By this, major weak spots can be eliminated in the Handle PID system and it becomes robust against core infrastructure outages, sudden network topology changes, packet loss and heavy load situations. The second goal of the thesis is the integration of next generation data dissemination technologies based on location-independent network paradigms into the domain of persistent identifier systems. Therefore, we propose to employ the Handle system for citing research datasets which are disseminated by location-independent technologies based on BitTorrent and NDN. To tackle the trust challenges of dynamic data locations, we create a novel approach for trusted data dissemination in location-independent networks that ensures the authenticity of data as well as the attribution to data issuers. This is done by incorporating the foundations of the Handle PID system and a further format for exchanging complex access information in PIDs.
- Research Article
23
- 10.1045/may2012-simons
- May 1, 2012
- D-Lib Magazine
Research is increasingly collaborative and global in nature, and efforts to manage the vast amounts of research data generated daily require global solutions. The Digital Object Identifier (DOI) system provides a means of persistent identification of research data collections and datasets that is global, standardised and widely used. The Australian National Data Service (ANDS) partnered with DataCite to offer a DOI minting service. At Griffith University, implementing DOIs raised governance questions common to other institutions that encouraged discussion and collaboration.
- Research Article
2
- 10.29379/jedem.v14i2.749
- Dec 23, 2022
- JeDEM - eJournal of eDemocracy and Open Government
Many public sector organisations (PSO) use SaaS solutions from dominant global providers. Implementation of these solutions may raise issues concerning both lawful data processing, and the obligations that those PSOs have to maintain their digital assets. One example is a large Swedish PSO which addressed these issues as part of the adoption and implementation of Microsoft 365. The study identifies challenges and presents an analysis of the organisational implementation of that SaaS solution, exposing legal issues that arose in that context. Findings show an absence of a documented risk analysis related to the PSO's use of that SaaS solution, covering data processing and maintenance of its digital assets. Recommendations are presented to facilitate a PSO's procurement and implementation of such a SaaS solution to address issues around data processing and the processing of digital assets.
- Research Article
- 10.59298/jciss/2025/111.293900
- Aug 10, 2025
- IDOSR JOURNAL OF CURRENT ISSUES IN SOCIAL SCIENCES
Federal government of Nigeria is revolutionizing its public service by introducing e-governance in its operations. This study therefore examined the effect of this new technology on the performance of the public sector in the Federal Capital Territory, Abuja. Specifically, the study determined the effect of e-governance on improved service delivery in public sector organizations in Abuja; investigated the effect of e-governance on improved transparency and accountability in public sector organizations in Abuja; and examined the challenges facing e-governance operations in public sector organizations in Abuja. Three research questions and three hypotheses guided the study. It was quantitative field survey research that made use of questionnaires. The study had a population of 1002 drawn from four public sector organizations in Abuja that have embraced e-governance in its operations. With Freud and Williams’s statistical sampling formula, this huge population was reduced to a researchable sampling size of 347. Independent t-test tool was used in the testing of the hypotheses leading to the findings. The study found that e-governance operations have a positive significant effect on service delivery and transparency and accountability in public sector organizations in Abuja. The study also found that challenges such as low digital literacy and infrastructure deficit inhibit e-governance operations in public sector organizations in Abuja. As a way out of this, the study recommended for a robust investment in ICT infrastructure by the government to expand its ICT infrastructure. It called also for capacity building of public servants so they can efficiently use e-governance tools and improve their service delivery functions more. Keywords: Technology, E-governance, Performance, Public Sector, Service Delivery, Transparency, Accountability
- Conference Article
106
- 10.1109/coinfo.2009.66
- Nov 1, 2009
Since 2005, the German National Library of Science and Technology (TIB) has offered a successful Digital Object Identifier (DOI) registration service for persistent identification of research data. In 2009, TIB, the British Library, the Library of the ETH Zurich, the French Institute for Scientific and Technical Information (INIST), the Technical Information Center of Denmark, Canada Institute for Scientific and Technical Information (CISTI) the Australien National Data Service (ANDS) and the Dutch TU Delft Library all signed a Memorandum of Understanding to improve access to research data on the internet. The goal of this cooperation is to establish a not-for-profit agency called DataCite that enables organisations to register research datasets and assign persistent identifiers to them, so that research datasets can be handled as independent, citable, unique scientific objects.
- Research Article
26
- 10.2139/ssrn.1639998
- Jan 1, 2010
- SSRN Electronic Journal
Since 2005, the German National Library of Science and Technology (TIB) has offered a successful Digital Object Identifier (DOI) registration service for persistent identification of research data. In 2009, TIB, the British Library, the Library of the ETH Zurich, the French Institute for Scientific and Technical Information (INIST), the Technical Information Center of Denmark, Canada Institute for Scientific and Technical Information (CISTI) the Australian National Data Service (ANDS) and the Dutch TU Delft Library all signed a Memorandum of Understanding to improve access to research data on the internet. The goal of this cooperation is to establish a not-for-profit agency called DataCite that enables organisations to register research datasets and assign persistent identifiers to them, so that research datasets can be handled as independent, citable, unique scientific objects.
- Book Chapter
3
- 10.1007/978-3-031-15086-9_6
- Jan 1, 2022
Lawful and appropriate use of cloud-based globally provided Software-as-a-Service (SaaS) solutions by a public sector organisation (PSO) for data processing and maintenance of digital assets presupposes an investigation of all relevant contract terms. Having obtained, analysed, and filed all relevant contract terms when using a SaaS solution is a prerequisite for good administration. Identifying and obtaining all relevant contract terms for a SaaS solution involves significant obstacles which in practice may be impossible to overcome for each PSO. This paper addresses how PSOs investigate contract terms prior to adoption, and why PSOs use a globally provided SaaS solution without having identified and obtained all relevant contract terms. Through a review of responses to questions and public documents from Swedish PSOs we analysed how each PSO had investigated contract terms and licences for the Microsoft 365 (M365) solution prior to adoption and use of the solution in each PSO. We find that no PSO had investigated all relevant contract terms prior to use of M365, which implies that each PSO uses M365 under unknown contract terms. Further, we find that all PSOs use M365 for data processing of its digital assets under unknown contract terms and that each PSO has significant dependence and trust in its supplier.
- Research Article
- 10.3389/frfst.2025.1658625
- Oct 9, 2025
- Frontiers in Food Science and Technology
The agri-livestock sector in Pakistan is of critical interest to the socioeconomic fabric of rural areas and is currently facing the risk of systemic vulnerabilities being exacerbated by climate change. The study is part of the doctoral thesis for performance evaluation of a public sector entity under environmental factors. In this paper, we will discuss adaptive measures toward climate-smart agriculture (CSA) in resilient livestock supply chains, based on a study of agri-livestock public sector organizations. In the study, we combine institutional records, interviews of key players, and thematic analysis research to extrapolate on the main facilitators and barriers to resilience. In the study, we conclude that structural gaps, such as unstructured logistics, lacking governance, insufficient financial inclusion, and gender discrimination, continue to exist; nevertheless, new innovations, such as silage production, solarized cool chain, and digital tracing systems, are developing. As much as interventions by public sector organizations can have local gains, resilience at the system level is limited by misaligning policies, institutional impediments, and inadequate investments in the climate-resistant infrastructure. In the paper, it is indicated that the adaptive capacity of the livestock sector in Pakistan requires a system-wide, coordinated, equity-based recalibration that needs to incorporate CSA principles in line with market access, public–private partnership (PPP), and inclusive governance. The policy recommendations underline the importance of integrated national policies, flexible investment in infrastructure, gender mainstreaming, and financial de-risking procedures to preserve food security and rural livelihoods when faced with intensifying climate risks.
- Research Article
- 10.59490/abe.2016.21.1525
- Jan 1, 2016
- Architecture and the Built Environment
If data are the building blocks to generate information needed to acquire knowledge and understanding, then geodata, i.e. data with a geographic component (geodata), are the building blocks for information vital for decision-making at all levels of government, for companies and for citizens. Governments collect geodata and create, develop and use geo-information - also referred to as spatial information - to carry out public tasks as almost all decision-making involves a geographic component, such as a location or demographic information. Geo-information is often considered “special” for technical, economic reasons and legal reasons. Geoinformation is considered special for technical reasons because geo-information is multi-dimensional, voluminous and often dynamic, and can be represented at multiple scales. Because of this complexity, geodata require specialised hardware, software, analysis tools and skills to collect, to process into information and to use geoinformation for analyses. Geo-information is considered special for economic reasons because of the economic aspects, which sets it apart from other products. The fixed production costs to create geo-information are high, especially for large-scale geo-information, such as topographic data, whereas the variable costs of reproduction are low which do not increase with the number of copies produced. In addition, there are substantial sunk costs, which cannot be recovered from the market. As such, geo-information shows characteristics of a public good, i.e. a good that is non-rivalrous and non-excludable. However, to protect the high investments costs, re-use of geo-information may be limited by legal and/or technological means such as intellectual property rights and digital rights management. Thus, by making geo-information excludable, it becomes a club good, i.e. a non-rivalrous but excludable good. By claiming intellectual property rights, such as copyright and/or database rights, and restricting (re-)use through licences and licence fees, geo-information can be commercially exploited and used to recover some of the investment costs. Geo-information is considered special for a number of legal reasons. First, as geo-information has a geographic component, e.g. a reference to a location, geoinformation may contain personal data, sensitive company data, environmentally sensitive data, or data that may pose a threat to the national security. Therefore, the dataset may have to be adapted, aggregated or anonymised before it can be made public. Secondly, geo-information may be subject to intellectual property rights. There may be a copyright on cartographic images or database rights on digital information. Such intellectual property rights may be claimed by third parties involved in the information chain, e.g. a private company supplying aerial photography to the National Mapping Authority. The data holder may also claim intellectual property rights to commercially exploit the dataset and recoup some of the vast investment costs made to produce the dataset. Lastly, there may be other (international) legislation or agreements that may either impede or promote publishing public sector information, whereby in some cases, these policies may contradict each other. It has been recognised that to deal with national, regional and global challenges, it is essential that geo-information collected by one level of government or government organisation be shared between all levels of government via a so-called Spatial Data Infrastructure (SDI). The main principles governing SDIs are that data are collected once and (re-)used many times; that data should be easy to discover, access and use; and that data are harmonised so that it is possible to combine spatial data from different sources seamlessly. In line with the SDI governing principles, this dissertation considers accessibility of information to include all these aspects. Accessibility concerns not only access to data, i.e. to be able to view the data without being able to alter the contents but also re-use of data, i.e. to be able to download and/or invoke the data and to share data, including to be able to provide feedback and/or to provide input for co-generated information. Accessibility to public sector geo-information is not only essential for effective and efficient government policy-making but is also associated with realising other ambitions. Examples of these ambitions are a more transparent and accountable government, more citizens’ participation in democratic processes, (co-)generation of solutions to societal problems, and to increase economic value due to companies creating innovative products and services with public sector information as a resource. Especially the latter ambition has been the subject of many international publications stressing the enormous potential economic value of re-use of public sector (geo-) information by companies. Previous research indicated that re-users of public sector information in Europe encountered barriers related to technical, organisational, legal and financial aspects, which was deemed to be the main reason why in Europe the number of value added products and services based on public service information were lagging compared to the United States. Especially the latter two barriers (restrictive licence conditions and high licence fees) were often cited to be the main barriers for reusers in Europe. However, in spite of considerable resources invested by governments to establish spatial data infrastructures, to facilitate data portals and to release public sector information as open data, i.e. without legal and financial restrictions, the expected surge of value added products based on public sector information has not quite eventuated to date and the expected benefits still appear to lag expectations. When this research started a decade ago, the debate around accessibility of public sector information focussed on access policies. Access policies ranged from open access (data available with a minimum of legal restrictions and for no more than marginal dissemination costs) to full cost recovery, whereby all costs incurred in collection, creation, processing, maintenance and dissemination costs to be recovered from the re-users. Most of the public sector bodies in the European Union adhered to a cost recovery policy for allowing re-use of public sector information. In 2003, the European Commission adopted two directives to ensure better accessibility of public sector information Directive 2003/4/EC of the European Parliament and of the Council of 28 January 2003 on public access to environmental information and repealing Council Directive 90/313/EEC, the so-called Access Directive, provided citizens the right of access to environmental information. Citizens should be able to access documents related to the environment via a register, preferably in an electronic form and if a copy of a document was requested, the charges must not exceed marginal dissemination costs. Directive 2003/98/EC of the European Parliament and of the Council of 17 November 2003 on the re-use of public sector information, the co-called PSI Directive, intended to create conditions for a level playing field for all re-users of public sector information. However, the PSI Directive of 2003 left room for public sector organisations to maintain a cost recovery regime with restrictive licence conditions. In spite of these directives, access policies for geographic data were slow to change in most European nations. At the end of the last decade, accessibility of public sector information received two major impulses. The first major impulse was the implementation of Directive 2007/2/ EC of the European Parliament and of the Council of 14 March 2007 establishing an Infrastructure for Spatial Information in the European Community (INSPIRE), the cocalled INSPIRE Directive, established a framework of standardisation rules for the data and publishing via web services, which significantly contributed to the accessibility of public sector geo-information. The second major impulse was the development of open data policies following the Digital Agenda for Europe adopted in 2010 and the USA Open Government Directive of 2009 and the Digital Agenda for Europe of 2010. These two impulses were the main drivers in Europe to start a careful move from cost recovery policies to open access or open data policies and for more public sector information to be made available as open data. Thus, of the four barriers to re-use of public sector information data cited in Chapter 1 (legal, financial, technical and organisational barriers), two barriers should have been lifted to a large degree due to open data. This shift to open data provided an excellent opportunity to test the hypothesis that the main barriers for re-users of public sector information were indeed restrictive licences and high fees as suggested by earlier research. Chapter 2 showed that by 2008, most European Union Member States had transposed and implemented the 2003/98/EC PSI Directive, however, in various ways and with considerable delay. By 2008, the effects of the PSI Directive were only slowly starting to emerge. A number of Member States reviewed their access policies and more public sector information became available for re-use. Some Member States made the information available free-of-charge or reduced their fees significantly. In many cases, where re-use fees were reduced the number of regular re-users increased significantly and total revenue even increased in spite of lower fees. Although the 2007/2/EC INSPIRE Directive paved the way for technical interoperability by providing guidelines for web services and catalogues, neither the INSPIRE Directive nor the PSI Directive had tackled the issue of legal interoperability. Chapter 2 also demonstrated that a major barrier to creating a level playing field for the private sector was the fact that some public sector bodies acted as value added resellers by developing and selling products and services based on their own data. Thus, the level playing field envisioned by the European Commission had not been realised. Chapter 3 researched the aspect of harmonised licences as a first step towards legal interoperability. Earlier research had indicated that one of the biggest barriers for re-users were complex, intransparent and inconsistent licence conditions, especially for re-users wanting to combine data from multiple sources. A survey of licences used by public sector data providers in the Netherlands demonstrated that although there were differences in length and language, there were also many similarities. The conclusion was that the introduction of a licence suite inspired by the Creative Commons concept would be a step towards increased transparency and consistency of geo-information license agreements. This chapter introduced a conceptual model for such a geo-information licence suite, the so-called Geo Shared licences. Both Creative Commons and Geo Shared licence suites enable harmonisation of licence conditions and promote transparency and legal interoperability, especially when re-users combine data from different sources. The Geo Shared licence suite became a serious option for inclusion into the draft version of the INSPIRE Directive as an annex. Unfortunately, the concept of one licence suite for the entire European Union came too early in 2006. The Geo Shared licences were further developed and implemented into the Dutch National Geo Register. In 2009, the European Commission recognised that PSI was the single largest source of information in Europe and the potential for re-use of PSI needed to be highlighted in the digital age. As part of a review of the 2003/98/EC PSI Directive, the European Commission carried out a round of consultations with stakeholders to seek their views on specific issues to be addressed in the future in 2010. In addition, the Commission commissioned a number of studies. These studies included a review of studies on public sector information re-use and related market studies, an assessment of the different models of supply and charging for public sector information and a study on public sector re-user in the cultural sector. The first study, carried out by Graham Vickery in 2011, showed that the overall economic gain from opening up public sector information as a resource for new products and services could be in the order of €40 billion per annum in the European Union. Both the Vickery Report and the second study, the so-called POPSIS Study, showed that for most public sector data providers their revenues from licence fees were relatively low in comparison to their total budget. After the evaluation, Directive 2013/37/EU of the European Parliament and of the Council of 26 June 2013 amending Directive 2003/98/EC on the re-use of public sector information was adopted and came into force on 17 July 2013. Chapter 4 described the main changes of the 2013/37/EU Amended PSI Directive, including the recommendation to employ open data licences. This chapter continued with a review of the various open data licences in use in Europe and analysed their interoperability. Although adoption of open data licences for public sector information should have addressed legal interoperability barriers for re-users, in practice, the different types of open data licences might not be so interoperable after all. Effectively, only a public domain declaration, such as a Creative Commons Zero (CC0) declaration, is suitable for open data re-users requiring with cross-border data sets and that such a public domain declaration is published in a prominent place to remove uncertainty for re-users. Without a public domain declaration, re-use of open data is still impeded as re-users are loathe to invest time into the development of value added products or services when it is uncertain if and which restrictions may be applicable and what the impact may be on their product or service. This dissertation also researched the financial and economic aspects of public sector information accessibility. Chapters 1 and 2 indicated that a cost recovery regime for dissemination of public sector information provided a financial barrier for private sector re-users because the fees charged were perceived to be too high. However, in 2008, there were still many advocates for maintaining a cost recovery regime. Especially public sector bodies that are not funded by the national Treasury, the socalled self-funding agencies, needed revenue from data sales to cover a substantial part of their operational costs. A sustainable source of revenue was viewed as essential to maintain the data at an adequate level, and to ensure actuality and continuity. Chapter 5 explored the potential business models and pricing mechanisms for public sector INSPIRE web services. Although, depending on the type of web service, and type of re-user, there might have been an argument for employing a subscription model as a pricing mechanism, business models based on generating revenue from public sector information would not be viable in the long run and were not in the spirit of the INSPIRE Directive. This research concluded that public sector information web services employing different pricing regimes were counterproductive to achieving financial interoperability. In Chapter 6, business models for public sector data providers were revisited, this time from an open data perspective. Government agencies, including self-funding government agencies are under increasing pressure to implement open data policies. This chapter analysed the business models of self-funding agencies either already providing open data or under pressure to provide (some) open data in the near future. The analysis showed which adaptions might be necessary to ensure the long-term availability of high quality open data and the long-term financial sustainability of self-funding agencies. The case studies confirmed that providing (raw) open data does not necessarily lead to losses in revenue in the long term as long as the organisation has enough flexibility to adapt its role in the information value chain, especially when revenue from licence fees represents only a relative small part of their total budget. The case studies indicated that switching to open data has resulted in internal efficiency gains. In practice, it is difficult to isolate and quantify the internal efficiency gains that are solely attributable to open data as the researched organisations continuously implement efficiency measures. However, the reported decreases in internal and external transaction costs due to open data are in line with the case study carried out in Chapter 7. Open data also provided an excellent opportunity to assess the effects of open data ex ante as baseline measurements could be carried out. To develop both quantitative and qualitative indicators to assess the success of a policy change is a challenge for open data initiatives. In Chapter 7, a model to assess the effects on the organisation of an open data provider was developed. Liander, a private energy network administrator mandated with a public task, planned to publish some of their datasets as open data in the autumn of 2013. This offered an excellent opportunity to apply the developed assessment model to provide an insight into internal, external, and relational effects on Liander. A benchmark was carried out prior to release of open data and a follow-up measurement one year later. The benchmark provided an insight into the then work processes and into the preparations required to implement open data. The follow-up monitor indicated that Liander open data are used by a wide range of users and have had a positive effect on the development of apps to aid energy savings. However, it remains a challenge to quantify the societal effects of such apps. The follow-up monitor also indicated that regular re-users of Liander data used the open data to improve existing applications and work processes rather than to create new products. The case study demonstrated that private energy companies could successfully release open data. The case study also showed that Liander served as a best-practice case for open data and had a flywheel effect on companies within the same sector. By 2015, nearly all energy network administrators had published similar open data. The monitoring model developed in this project was assessed to be suitable to monitor the open data effects on the organisation of the data provider. The assessment model developed and tested in Chapter 7 proved to be suitable to monitor the effects of open data on organisational level. However, to provide a more complete picture of the effects of open data and to assess if there are other barriers for re-users, a more holistic approach was required to assess the maturity of open data. Therefore, a holistic open data assessment framework addressing the supplier side, the governance side, and the user side of the open data was developed and applied to the Dutch open data infrastructure in Chapter 8. This Holistic Open Data Maturity Assessment Framework was used to evaluate the State of the Open Data Nation in the Netherlands and to provide valuable information on (potential) bottlenecks. The framework showed that geographic data scored significantly better than other types of government data. The standardisation and implementation rules laid down by INSPIRE Directive framework appear to have been a catalyst for moving geographic data to a higher level of maturity. The maturity assessment framework provided Dutch policy makers with useful inputs for further development of the open data ecosystem and development of well-founded strategies that will ensure the full potential of open data will be reached. Since the publication of the State of the Open Data Nation in 2014, a number of the recommendations have already been implemented. This dissertation demonstrated that many aspects that should facilitate accessibility, such as standardised metadata, have already been addressed for geodata. This research also showed that for other types of data, there is still a long way to go. There is a growing demand for other types of data, such as financial data and healthcare data. Public sector organisations holding such types of data need hands-on guidelines to enable publication of their datasets, preferably as open data. However, data published as open data are forever and cannot be recalled. Therefore, the decision to publish public sector data as open data is complex: datasets are often of a heterogeneous nature and may contain microdata (data that quantify observations or facts, such as data collected during surveys) Although microdata may not necessarily contain personal data, the datasets will probably have to be processed before publication to address confidentiality and data quality issues. In addition, there is a tension between open data and protection of personal data. The big question remains to which level the data need to be aggregated and/or anonymised to ensure protection of personal data now and in the future, and at the same time keeping sufficient significance to be re-usable. Another issue that needs further research is data-ownership of sensor data and co-created data. Increasingly, sensor data generated by e.g. smart phones, smart energy meters and traffic sensors are collected by the public sector and the private sector and become part of a big data ecosystem. In addition, public sector organisations cooperate with other public sector organisations and the private sector to create information from their data, so-called co-created information. Citizens also collect data or complement information on a voluntary basis, e.g. bird counts data. Co-created information will become more commonplace in the coming decades, as will the contribution of sensor data to a big data ecosystem. However, the aspect of who owns the data in which part of the information value chain has not been researched. Uncertainty related to third party rights will pose a barrier to publishing open data. Therefore, the aspect of data-ownership for sensor data and for co-created data should be further researched.
- Research Article
5
- 10.3390/systems11080424
- Aug 13, 2023
- Systems
Efficient monitoring and achievement of the Sustainable Development Goals (SDGs) has increased the need for a variety of data and statistics. The massive increase in data gathering through social networks, traditional business systems, and Internet of Things (IoT)-based sensor devices raises real questions regarding the capacity of national statistical systems (NSS) for utilizing big data sources. Further, in this current era, big data is captured through sensor-based systems in public sector organizations. To gauge the capacity of public sector institutions in this regard, this work provides an indicator to monitor the processing capacity of the public sector organizations within the country (Pakistan). Some of the indicators related to measuring the capacity of the NSS were captured through a census-based survey. At the same time, convex logistic principal component analysis was used to develop scores and relative capacity indicators. The findings show that most organizations hesitate to disseminate data due to concerns about data privacy and that public sector organizations’ IT personnel are unable to deal with big data sources to generate official statistics. Artificial intelligence (AI) techniques can be used to overcome these challenges, such as automating data processing, improving data privacy and security, and enhancing the capabilities of IT human resources. This research helps to design capacity-building initiatives for public sector organizations in weak dimensions, focusing on leveraging AI to enhance the production of quality and reliable statistics.
- Preprint Article
- 10.5194/egusphere-egu23-1599
- May 15, 2023
The consequences of global change on the ocean are multiple such as increase in temperature and sea level, stronger storms, deoxygenation, impacts on ecosystems. But the detection of changes and impacts is still difficult because of the diversity and variability of marine environments. While there has been a clear increase in the number of marine and coastal observations, whether by in situ, laboratory or remote sensing measurements, each data is both costly to acquire and unique. The number and variety of data acquisition techniques require efficient methods of improving data availability via interoperable portals, which facilitate data sharing according to FAIR principles for producers and users. ODATIS, the ocean cluster of Data Terra, the French research infrastructure for Earth data, is the entry point to access all the French Ocean observation data (Ocean Data Information and Services ; www.odatis-ocean.fr/en/). The first challenge of ODATIS is to get data producers to share data. To that purpose, ODATIS offers several services to help them define Data Management Plan (DPM), implement the FAIR principles, make data more visible and accessible by being referenced in the ODATIS catalog, and better tracked and cited through a Digital Object Identifier (DOI). ODATIS also offers a service for publishing open scientific data on the sea, through SEANOE (www.seanoe.org) that provides a DOI that can be cited in scientific articles in a reliable and sustainable way. In parallel to the informatic development of the ocean cluster, further communication and training are needed to inform the research community of these new tools. Through technical workshops, Odatis offers data providers practical experience and support in implementing data access, visualization and processing services. Finally, ODATIS relies on scientific consortia in order to promote and develop innovative processing methods and products for remote, airborne, or in situ observations of the ocean and its interfaces (atmosphere, coastline, seafloor) with the other clusters of the RI Data Terra.