Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Lost in LOD: Analyzing the Linked Open Data Cloud Quality Maze

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Since 2007, the Linked Open Data (LOD) Cloud has served as a central hub for datasets following Linked Data (LD) principles, offering a large repository of interconnected information. Over time, it has undergone multiple quality assessments to ensure datasets are accessible, well-maintained, and meet standards. Grounded on metadata assessment performed over time, this paper examines the current quality of the LOD Cloud by analyzing 1,658 datasets from the December 2024 snapshot, evaluated against 52 quality metrics. By proposing a reproducible methodology, it reports about the quality assessment and the trend analysis assessing progress, identifying persistent problems, and verifying how datasets registered in the LOD Cloud evolve over time. According to results, many earlier issues persist. Datasets still lack consistency in metadata structure, licenses, and distribution format. Moreover, they mainly remain in archived versions, with real-time access often poorly maintained. Well-curated, up-to-date datasets are exceptions rather than the rule.

Similar Papers
  • Book Chapter
  • Cite Count Icon 3
  • 10.1007/978-3-642-45263-5_4
LiQuate-Estimating the Quality of Links in the Linking Open Data Cloud
  • Jan 1, 2013
  • Edna Ruckhaus + 1 more

During the last years, RDF datasets from almost any knowledge domain have been published in the Linking Open Data (LOD) cloud. The Linked Open Data guidelines establish the conditions to be satisfied by resources in order to be included as part of the LOD cloud, as well as connected to previously published data. The process of publication and linkage of resources in the LOD cloud relies on: i) data cleaning and transformation into existing RDF formats, ii) storage of the data into RDF storage systems, and iii) data interlinking. Because of data source heterogeneity, generated RDF data may be ambiguous and links may be incomplete with respect to this data. Users of the Web of Data require linked data to meet high quality standards in order to develop applications that can produce trustworthy results, but data in the LOD cloud has not been curated; thus, tools are necessary to detect data quality problems. For example, researchers that study Life Sciences datasets to explain phenomena or identify anomalies, demand that their findings correspond to current discoveries, and not to the effect of low data quality standards of completeness or redundancy. In this paper we propose LiQuate, a system that uses Bayesian networks to study the incompleteness of links, and ambiguities between labels and between links in the LOD cloud, and can be applied to any domain. Additionally, a probabilistic rule-based system is used to infer new links that associate equivalent resources, and allow to resolve the ambiguities and incompleteness identified during the exploration of the Bayesian network. As a proof of concept, we applied LiQuate to existing Life Sciences linked datasets, and detected ambiguities in the data, that may compromise the confidence of the results of applications such as link prediction or pattern discovery. We illustrate a variety of identified problems and propose a set of enriched intra- and inter-links that may improve the quality of data items and links of specific datasets of the LOD cloud.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/icsc.2014.27
Analyzing the Applicability of the Linking Open Data Cloud for Context-Aware Services
  • Jun 1, 2014
  • Moritz Von Hoffen + 2 more

The amount of data within the Linking Open Data (LOD) cloud is steadily increasing and resembles a rich source of information. Since Context-aware Services (CAS) can highly benefit from background information, e.g., about the environment of a user, it makes sense to leverage that enormous amount of data already present in the LOD cloud to enhance the quality of these services. Within this work, the applicability of the LOD cloud as provider for contextual information to enrich CAS is investigated. For this purpose, non-functional criteria of discoverability and availability are analyzed, followed by a presentation of an overview of the different domains covered by the LOD cloud. In order to ease the process of finding a dataset that matches the information needs of a developer of a CAS, techniques for retrieving contents of LOD datasets are discussed and different approaches to condense the dataset to its most important concepts are shown.

  • Research Article
  • Cite Count Icon 1
  • 10.1142/s1793351x14400121
Linked Open Data for Context-aware Services: Analysis, Classification and Context Data Discovery
  • Dec 1, 2014
  • International Journal of Semantic Computing
  • Moritz Von Hoffen + 1 more

The amount of data within the Linking Open Data (LOD) Cloud is steadily increasing and resembles a rich source of information. Since Context-aware Services (CAS) are based on the correlation of heterogeneous data sources for deriving the contextual situation of a target, it makes sense to leverage that enormous amount of data already present in the LOD Cloud to enhance the quality of these services. Within this work, the applicability of the LOD Cloud as a context provider for enriching CAS is investigated. For this purpose, a deep analysis according to the discoverability and availability of datasets is performed. Furthermore, in order to ease the process of finding a dataset that matches the information needs of a CAS developer, techniques for retrieving contents of LOD datasets are discussed and different approaches to condense the dataset to its most important concepts are shown. Finally, a Context Data Lookup Service is introduced that enables context data discovery within the LOD Cloud and its applicability is highlighted based on an example.

  • Conference Article
  • Cite Count Icon 4
  • 10.1109/imis.2014.95
Keyword-Based SPARQL Query Generation System to Improve Semantic Tractability on LOD Cloud
  • Jul 1, 2014
  • Soyeon Im + 3 more

As the number of RDF triples on the Linking Open Data (LOD) Cloud has been exponentially increased, the difficulties of information query have been increased. To query information on the LOD Cloud, the users have to have some capabilities for developing SPARQL and/or RDQL query statements and exact knowledge for web resources with RDF such as URI, DB title, and name of things. However, it is almost impossible to require the capabilities and/or knowledge to the users who are familiar with keywords search. So, we propose fully automated keyword-based SPARQL query generation system. In our system, the users can query information on the LOD Cloud without having capabilities about structured query language and having prior knowledge of web resources with RDF. The users should type a set of keywords into our system. To do so, we developed a propertybased path finding algorithm and an automated SPARQL query generation that can be used to provide query recommendations for the users. In the experimental section, we illustrate an example to validate the effectiveness of our system, and perform a simulation to show the superiority of the algorithms. The experiment results are not bad for a first attempt. If we improve the algorithms, we can expect a better result.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/icosc.2015.7050776
Linked crowdsourced data - Enabling location analytics in the linking open data cloud
  • Feb 1, 2015
  • Abdulbaki Uzun

Geospatial datasets in the Linking Open Data (LOD) Cloud are rather of static nature and mainly consist of information such as a name, geo coordinates, an address, or opening hours. There is no linked dataset providing dynamic information about the "popularity" of certain places or the "Visiting frequency" of users in specific contextual situations. This type of information within the LOD Cloud, however, would enable a variety of new applications based on semantically enriched location analytics. In this paper, we present Linked Crowdsourced Data as a dataset, which links real user location preferences (e.g., check-ins, ratings, or comments) as well as specific context situations (e.g., weather conditions, holiday information, or measured networks) collected via crowdsourcing to static location data. We showcase the applicability of this dataset for location analytics use cases through a map visualization and highlight its added value with exemplary SPARQL queries that allow for location requests depending on historic context information.

  • Book Chapter
  • Cite Count Icon 1
  • 10.1007/978-3-030-36599-8_23
OntoPPI: Towards Data Formalization on the Prediction of Protein Interactions
  • Jan 1, 2019
  • Yasmmin Cortes Martins + 4 more

The Linking Open Data (LOD) cloud is a global data space for publishing and linking structured data on the Web. The idea is to facilitate the integration, exchange, and processing of data. The LOD cloud already includes a lot of datasets that are related to the biological area. Nevertheless, most of the datasets about protein interactions do not use metadata standards. This means that they do not follow the LOD requirements and, consequently, hamper data integration. This problem has impacts on the information retrieval, specially with respect to datasets provenance and reuse in further prediction experiments. This paper proposes an ontology to describe and unite the four main kinds of data in a single prediction experiment environment: (i) information about the experiment itself; (ii) description and reference to the datasets used in an experiment; (iii) information about each protein involved in the candidate pairs. They correspond to the biological information that describes them and normally involves integration with other datasets; and, finally, (iv) information about the prediction scores organized by evidence and the final prediction. Additionally, we also present some case studies that illustrate the relevance of our proposal, by showing how queries can retrieve useful information.

  • Book Chapter
  • Cite Count Icon 2
  • 10.1007/978-3-662-45550-0_48
Traversing the Linking Open Data Cloud to Create News from Tweets
  • Jan 1, 2014
  • Francisco Berrizbeita + 1 more

We propose a two-fold approach that is able to both consume and exploit semantics encoded in the Linking Open Data (LOD) cloud, and create news that document events reported in micro-blogging posts that correspond to documentary tweets. A documentary tweet is similar to a newspaper headline and reports an incident or event. Knowledge extracted from documentary tweets are used to develop a story line which will be augmented with RDF facts consumed from the LOD cloud. The resulting news content is represented in RDF using the rNews Ontology, facilitating news generation and retrieval. We study effectiveness of our approach with respect to a gold standard of manually tagged tweets. Initial experimental results suggest that our techniques are able to generate content that reflects up to 76.38% of the manually tagged terms.KeywordsNews ArticleSentiment AnalysisStory LineNewspaper HeadlineNews ContentThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

  • Research Article
  • Cite Count Icon 2
  • 10.4314/stech.v7i1.3
Making attributes from the Linked Open Data (LOD) cloud a part of Spatial Data Infrastructures (SDIS) for thematic mapping
  • Apr 19, 2018
  • AFRREV STECH: An International Journal of Science and Technology
  • Wiafe Owusu-Banahene

The creation of the linked open data (LOD) cloud has enhanced the availability of interlinked open statistical data associated with geographic regions (attributes) through the Web. Spatial data infrastructures (SDIs) are widely used to share, discover, visualise and retrieve geospatial data. Geoportals are the visible parts of SDIs focused on interoperability through implementing standards such OGC web services for discovery and use of geographic data and services. In the current form OGC services cannot be directly connected to the LOD cloud. This research aimed at finding novel ways of visualising linked data in the form of thematic maps over the Internet. This paper provided answers to the question of how an OGC Web Map Service (WMS) can create thematic maps by combining attributes from the LOD cloud with geometry stored in a spatial database server (SDS). This research contributed to bridging the gap between linked data, SDI and web thematic maps and further showed how existing web mapping and OGC technologies can benefit from the Semantic Web. First, the design of a geospatial web service (representing the visible part of an SDI) that accesses attribute data from the LOD cloud referred to in this paper as SDI-LOD is presented. SDI-LOD produces web thematic maps by programmatically combining attributes from the LOD cloud with geometry in an SDS. Next, the author presented implementation results of SDI-LOD and concluded with a discussion of the implementation. This paper has motivated future work on SDI-LOD to integrating data from internet of things.Keywords: thematic map; LOD cloud; geospatial data; linked open data; geospatial web service; SDI

  • Conference Article
  • 10.1145/3019612.3019933
Discovering and linking with life sciences linked open data cloud
  • Apr 3, 2017
  • Muntazir Mehdi

Linked Open Data (LOD) Cloud is a mesh of open datasets coming from different domains. Among these datasets, a no-table amount of datasets belong to the life sciences domain and are linked together forming an interlinked Life Sciences Linked Open Data (LSLOD) Cloud. Link generation within LOD cloud is a hectic job for publishers. Addressing this challenge, frameworks like Silk 1 and LIMES 2 have been proposed to help publishers link their local datasets to a remote LOD dataset. However, these linking framework rely on explicit information about datasets to link with. Currently, identification of pertinent datasets relies on manual effort and requires experience and knowledge of datasets. One potential method is to manually inspect the list of datasets, however, only high level and limited dataset descriptions are available regarding the topic of the dataset and available access mechanisms [10].

  • Conference Article
  • Cite Count Icon 19
  • 10.1145/2872518.2890545
LODVader
  • Jan 1, 2016
  • Ciro Baron Neto + 4 more

The Linked Open Data (LOD) cloud is in danger of becoming a black box. Simple questions such as "What kind of datasets are in the LOD cloud?", "In what way(s) are these datasets connected?" -- albeit frequently asked -- are at the moment still difficult to answer due to the lack of proper tooling support. The infrequent update of the static LOD cloud diagram adds to the current dilemma, since there is neither reliable nor timely-updated information to perform an interactive search, analysis or in particular visualization in order to gain insight into the current state of Linked Open Data. In this paper, we propose a new hybrid system which combines LOD Visualisation, Analytics and DiscovERy (LODVader) to aid in answering the above questions. LODVader is equipped with (1) a multi-layer LOD cloud visualization component comprising datasets, subsets and vocabularies, (2) dataset analysis components that extend the state of the art with new similarity measures and efficient link extracting techniques and (3) a fast search index that is an entry point for dataset discovery. At its core, LODVader employs a timely-updated index using a complex cluster of Bloom filters as a fast search index with low memory footprint. This BF cluster is able to efficiently perform analysis on link and dataset similarities based on stored predicate and object information, which -- once inverted -- can be employed to discover invalid links by displaying the Dark LOD Cloud. By combining all these features, we allow for an up-to-date, multi-dimensional LOD cloud analysis, which -- to the best of our knowledge -- was not possible before.

  • Book Chapter
  • Cite Count Icon 19
  • 10.1007/978-3-642-41033-8_80
Analyzing Linked Data Quality with LiQuate
  • Jan 1, 2013
  • Edna Ruckhaus + 2 more

In the last years, the number of datasets in the Linking Open Data (LOD) cloud and the applications that rely on links between these datasets to discover patterns or potential new associations, have exploded. However, because of data source heterogeneity, published data may suffer of redundancy, inconsistencies or may be incomplete; thus, results generated by linked data based applications may be imprecise or unreliable. We illustrate LiQuate (Linked Data Quality Assessment), a tool that combines Bayesian Networks and rule-based systems to analyze the quality of data and links in the LOD cloud.

  • Book Chapter
  • Cite Count Icon 21
  • 10.1007/978-3-319-11955-7_72
Analyzing Linked Data Quality with LiQuate
  • Jan 1, 2014
  • Edna Ruckhaus + 4 more

The number of datasets in the Linking Open Data (LOD) cloud as well as LOD-based applications have exploded in the last years. However, because of data source heterogeneity, published data may suffer of redundancy, inconsistencies, or may be incomplete; thus, results generated by LOD-based applications may be imprecise, ambiguous, or unreliable. We demonstrate the capabilities of LiQuate (Linked Data Quality Assessment), a tool that relies on Bayesian Networks to analyze the quality of data and links in the LOD cloud.

  • Book Chapter
  • Cite Count Icon 4
  • 10.1007/978-3-642-35085-6_12
Building a Linked Open Data Cloud of Linguistic Resources: Motivations and Developments
  • Jan 1, 2013
  • Christian Chiarcos + 4 more

We describe on going community-efforts to create a Linked Open Data (sub-)cloud of linguistic resources, with an emphasis on resources that are specific to linguistic research, namely annotated corpora and linguistic databases. We argue that for both types of resources, the application of the Linked Open Data paradigm and the representation in RDF represents a promising approach to address interoperability problems, and to integrate information from different repositories. This is illustrated with example studies for different kinds of linguistic resources.The efforts described in this chapter are conducted in the context of the Open Linguistics Working Group (OWLG) of the Open Knowledge Foundation. The OWLG is a network of researchers interested in linguistic resources and/or their publication under open licenses, and a number of its members are engaged in the application of the Linked Open Data paradigm to their resources. Under the umbrella of the OWLG, these efforts will eventually emerge in the creation of a Linguistic Linked Open Data cloud (LLOD).

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 33
  • 10.7717/peerj-cs.37
Semantic representation of scientific literature: bringing claims, contributions and named entities onto the Linked Open Data cloud
  • Dec 9, 2015
  • PeerJ Computer Science
  • Bahar Sateli + 1 more

Motivation.Finding relevant scientific literature is one of the essential tasks researchers are facing on a daily basis. Digital libraries and web information retrieval techniques provide rapid access to a vast amount of scientific literature. However, no further automated support is available that would enable fine-grained access to the knowledge ‘stored’ in these documents. The emerging domain ofSemantic Publishingaims at making scientific knowledge accessible to both humans and machines, by adding semantic annotations to content, such as a publication’s contributions, methods, or application domains. However, despite the promises of better knowledge access, the manual annotation of existing research literature is prohibitively expensive for wide-spread adoption. We argue that a novel combination of three distinct methods can significantly advance this vision in a fully-automated way: (i) Natural Language Processing (NLP) forRhetorical Entity(RE) detection; (ii)Named Entity(NE) recognition based on the Linked Open Data (LOD) cloud; and (iii) automatic knowledge base construction for both NEs and REs using semantic web ontologies that interconnect entities in documents with the machine-readable LOD cloud.Results.We present a complete workflow to transform scientific literature into a semantic knowledge base, based on the W3C standards RDF and RDFS. A text mining pipeline, implemented based on the GATE framework, automatically extracts rhetorical entities of typeClaimsandContributionsfrom full-text scientific literature. These REs are further enriched with named entities, represented as URIs to the linked open data cloud, by integrating the DBpedia Spotlight tool into our workflow. Text mining results are stored in a knowledge base through a flexible export process that provides for a dynamic mapping of semantic annotations to LOD vocabularies through rules stored in the knowledge base. We created a gold standard corpus from computer science conference proceedings and journal articles, whereClaimandContributionsentences are manually annotated with their respective types using LOD URIs. The performance of the RE detection phase is evaluated against this corpus, where it achieves an averageF-measure of 0.73. We further demonstrate a number of semantic queries that show how the generated knowledge base can provide support for numerous use cases in managing scientific literature.Availability.All software presented in this paper is available under open source licenses athttp://www.semanticsoftware.info/semantic-scientific-literature-peerj-2015-supplements. Development releases of individual components are additionally available on our GitHub page athttps://github.com/SemanticSoftwareLab.

  • Single Book
  • Cite Count Icon 2
  • 10.5771/9783956506611
Linking Knowledge
  • Jan 1, 2021
  • Richard P Smiraglia

The growth and population of the Semantic Web, especially the Linked Open Data (LOD) Cloud, has brought to the fore the challenges of ordering knowledge for data mining on an unprecedented scale. The LOD Cloud is structured from billions of elements of knowledge and pointers to knowledge organization systems (KOSs) such as ontologies, taxonomies, typologies, thesauri, etc. The variant and heterogeneous knowledge areas that comprise the social sciences and humanities (SSH), including cultural heritage applications are bringing multi-dimensional richness to the LOD Cloud. Each such application arrives with its own challenges regarding KOSs in the Cloud. With contributions by Sören Auer, Gerard Coen, Kathleen Gregory, Mohamad Yaser Jaradeh, Daniel Martínez Ávila, Philipp Mayr, Allard Oelen, Cristina Pattuelli, Tobias Renwick, Andrea Scharnhorst, Ronald Siebes, Aida Slavic, Richard P Smiraglia, Markus Stocker, Rick Szostak, Marnix van Berchum, Charles van den Heuvel, J. Bradford Young, Veruska Zamborlini and Marcia Zeng.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant