Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

From unstructured text to structured reasoning: a hybrid knowledge graph for Indonesian sentencing analysis

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

From unstructured text to structured reasoning: a hybrid knowledge graph for Indonesian sentencing analysis

Similar Papers
  • Research Article
  • Cite Count Icon 3
  • 10.1007/978-1-0716-3449-3_10
Natural Language Processing for Drug Discovery Knowledge Graphs: Promises and Pitfalls.
  • Sep 14, 2023
  • Methods in molecular biology (Clifton, N.J.)
  • J Charles G Jeynes + 2 more

Building and analyzing knowledge graphs (KGs) to aid drug discovery is a topical area of research. A salient feature of KGs is their ability to combine many heterogeneous data sources in a format that facilitates discovering connections. The utility of KGs has been exemplified in areas such as drug repurposing, with insights made through manual exploration and modeling of the data. In this chapter, we discuss promises and pitfalls of using natural language processing (NLP) to mine "unstructured text"- typically from scientific literature- as a data source for KGs. This draws on our experience of initially parsing "structured" data sources-such as ChEMBL-as the basis for data within a KG, and then enriching or expanding upon them using NLP. The fundamental promise of NLP for KGs is the automated extraction of data from millions of documents-a task practically impossible to do via human curation alone. However, there are many potential pitfalls in NLP-KG pipelines, such as incorrect named entity recognition and ontology linking, all of which could ultimately lead to erroneous inferences and conclusions.

  • Research Article
  • Cite Count Icon 31
  • 10.1080/13658816.2014.989855
Thematic signatures for cleansing and enriching place-related linked data
  • Mar 3, 2015
  • International Journal of Geographical Information Science
  • Benjamin Adams + 1 more

There has been significant progress transforming semi-structured data about places into knowledge graphs that can be used in a wide variety of geographic information systems such as digital gazetteers or geographic information retrieval systems. For instance, in addition to information about events, actors, and objects, DBpedia contains data about hundreds of thousands of places from Wikipedia and publishes it as Linked Data. Repositories that store data about places are among the most interlinked hubs on the Linked Data cloud. However, most content about places resides in unstructured natural language text, and therefore it is not captured in these knowledge graphs. Instead, place representations are limited to facts such as their population counts, geographic locations, and relations to other entities, for example, headquarters of companies or historical figures. In this paper, we present a novel method to enrich the information stored about places in knowledge graphs using thematic signatures that are derived from unstructured text through the process of topic modeling. As proof of concept, we demonstrate that this enables the automatic categorization of articles into place types defined in the DBpedia ontology (e.g., mountain) and also provides a mechanism to infer relationships between place types that are not captured in existing ontologies. This method can also be used to uncover miscategorized places, which is a common problem arising from the automatic lifting of unstructured and semi-structured data.

  • Research Article
  • Cite Count Icon 3
  • 10.1111/exsy.13617
Doc‐KG: Unstructured documents to knowledge graph construction, identification and validation with Wikidata
  • May 8, 2024
  • Expert Systems
  • Muhammad Salman + 3 more

The exponential growth of textual data in the digital era underlines the pivotal role of Knowledge Graphs (KGs) in effectively storing, managing, and utilizing this vast reservoir of information. Despite the copious amounts of text available on the web, a significant portion remains unstructured, presenting a substantial barrier to the automatic construction and enrichment of KGs. To address this issue, we introduce an enhanced Doc‐KG model, a sophisticated approach designed to transform unstructured documents into structured knowledge by generating local KGs and mapping these to a target KG, such as Wikidata. Our model innovatively leverages syntactic information to extract entities and predicates efficiently, integrating them into triples with improved accuracy. Furthermore, the Doc‐KG model's performance surpasses existing methodologies by utilizing advanced algorithms for both the extraction of triples and their subsequent identification within Wikidata, employing Wikidata's Unified Resource Identifiers for precise mapping. This dual capability not only facilitates the construction of KGs directly from unstructured texts but also enhances the process of identifying triple mentions within Wikidata, marking a significant advancement in the domain. Our comprehensive evaluation, conducted using the renowned WebNLG benchmark dataset, reveals the Doc‐KG model's superior performance in triple extraction tasks, achieving an unprecedented accuracy rate of 86.64%. In the domain of triple identification, the model demonstrated exceptional efficacy by mapping 61.35% of the local KG to Wikidata, thereby contributing 38.65% of novel information for KG enrichment. A qualitative analysis based on a manually annotated dataset further confirms the model's excellence, outshining baseline methods in extracting high‐fidelity triples. This research embodies a novel contribution to the field of knowledge extraction and management, offering a robust framework for the semantic structuring of unstructured data and paving the way for the next generation of KGs.

  • PDF Download Icon
  • Conference Article
  • Cite Count Icon 16
  • 10.18653/v1/d15-1061
An Entity-centric Approach for Overcoming Knowledge Graph Sparsity
  • Jan 1, 2015
  • Manjunath Hegde + 1 more

Automatic construction of knowledge graphs (KGs) from unstructured text has received considerable attention in recent research, resulting in the construction of several KGs with millions of entities (nodes) and facts (edges) among them.Unfortunately, such KGs tend to be severely sparse in terms of number of facts known for a given entity, i.e., have low knowledge density.For example, the NELL KG consists of only 1.34 facts per entity.Unfortunately, such low knowledge density makes it challenging to use such KGs in real-world applications.In contrast to best-effort extraction paradigms followed in the construction of such KGs, in this paper we argue in favor of ENTIty Centric Expansion (ENTICE), an entity-centric KG population framework, to alleviate the low knowledge density problem in existing KGs.By using ENTICE, we are able to increase NELL's knowledge density by a factor of 7.7 at 75.5% accuracy.Additionally, we are also able to extend the ontology discovering new relations and entities.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 5
  • 10.1007/s40747-022-00805-7
Exploiting lexical patterns for knowledge graph construction from unstructured text in Spanish
  • Aug 25, 2022
  • Complex & Intelligent Systems
  • Ana B Rios-Alvarado + 5 more

Knowledge graphs (KGs) are useful data structures for the integration, retrieval, dissemination, and inference of information in various information domains. One of the main challenges in building KGs is the extraction of named entities (nodes) and their relations (edges), particularly when processing unstructured text as it has no semantic descriptions. Generating KGs from texts written in Spanish represents a research challenge as the existing structures, models, and strategies designed for other languages are not compatible in this scenario. This paper proposes a method to design and construct KGs from unstructured text in Spanish. We defined lexical patterns to extract named entities and (non) taxonomic, equivalence, and composition relations. Next, named entities are linked and enriched with DBpedia resources through a strategy based on SPARQL queries. Finally, OWL properties are defined from the predicate relations for creating resource description framework (RDF) triples. We evaluated the performance of the proposed method to determine the degree of elements extracted from the input text and to assess their quality through standard information retrieval measures. The evaluation revealed the feasibility of the proposed method to extract RDF triples from datasets in general and computer science domains. Competitive results were observed by comparing our method regarding an existing approach from the literature.

  • Research Article
  • Cite Count Icon 5
  • 10.1142/s0219686725500180
An Intelligent Framework of Equipment Fault Diagnosis Based on Knowledge Graph
  • Oct 28, 2024
  • Journal of Advanced Manufacturing Systems
  • Shichen Zhai + 4 more

Equipment fault information typically exhibits the characteristics of fragmentation and diversified structure. The existing fault diagnosis methods are incapable of fully exploiting the prior knowledge and expert knowledge within the field, and the diagnosis results are overly one-sided. Given that it is challenging to obtain effective fault diagnostic knowledge for complex equipment and complex faults, this paper proposes the application of the domain knowledge graph (KG) for fault diagnosis. Different from the existing fault diagnosis methods involving the KG, our framework consists of two major parts. The first part is to construct an equipment fault KG from semi-structured and unstructured text. We further enrich the graph knowledge through knowledge completion, which can furnish high-quality knowledge sources for downstream applications. The second part is to employ the built fault KG either online or offline for fault diagnosis. We offer two approaches: the deep learning plus KG approach; the question-answering approach. The former not only guarantees the diagnostic accuracy but also provides more comprehensive diagnostic information. This constitutes the online utilization of the KG. The latter realizes the offline use of the KG, providing users with a natural and user-friendly manner to retrieve fault diagnosis information. We demonstrate and verify our proposed framework in the context of the bearing fault diagnosis.

  • Research Article
  • Cite Count Icon 14
  • 10.1016/j.displa.2024.102820
Knowledge graph of agricultural engineering technology based on large language model
  • Sep 26, 2024
  • Displays
  • Haowen Wang + 1 more

Knowledge graph of agricultural engineering technology based on large language model

  • Research Article
  • Cite Count Icon 8
  • 10.4018/jdm.2021100104
Low-Quality Error Detection for Noisy Knowledge Graphs
  • Oct 1, 2021
  • Journal of Database Management
  • *Chenyang Bu + 3 more

The automatic construction of knowledge graphs (KGs) from multiple data sources has received increasing attention. The automatic construction process inevitably brings considerable noise, especially in the construction of KGs from unstructured text. The noise in a KG can be divided into two categories: factual noise and low-quality noise. Factual noise refers to plausible triples that meet the requirements of ontology constraints. For example, the plausible triple <New_York, IsCapitalOf, America> satisfies the constraints that the head entity “New_York” is a city and the tail entity “America” belongs to a country. Low-quality noise denotes the obvious errors commonly created in information extraction processes. This study focuses on entity type errors. Most existing approaches concentrate on refining an existing KG, assuming that the type information of most entities or the ontology information in the KG is known in advance. However, such methods may not be suitable at the start of a KG's construction. Therefore, the authors propose an effective framework to eliminate entity type errors. The experimental results demonstrate the effectiveness of the proposed method.

  • Research Article
  • Cite Count Icon 6
  • 10.1109/access.2024.3462635
Knowledge Graph Generation and Application for Unstructured Data Using Data Processing Pipeline
  • Jan 1, 2024
  • IEEE Access
  • Sushmi Thushara Sukumar + 3 more

With the rapid advancement of technology and the vast volume of unstructured data available on the Internet, there is a pressing need to extract information from diverse data formats effectively. This is essential as valuable pieces of information may be lost. To address this issue, researchers are using Machine Learning (ML) and Natural Language Processing (NLP) techniques to extract information from unstructured text, including the utilization of Knowledge Graphs (KGs). This paper demonstrates end-to-end experimental studies of KG construction from unstructured text using open-source techniques and concrete real-world examples in different problem domains. The unstructured data underwent a text processing pipeline consisting of coreference resolution, named entity linking, and relationship extraction. The pipeline is designed to support automatic data storage in a graph database known as Neo4j. This storage includes the extracted entities and their relationships. Experiments were conducted on a real-world unstructured BBC News Dataset to analyze the outcome obtained from the pipeline. The experience can facilitate the adoption of KG creation for practitioners to capture valuable information from a large volume of unstructured text. The results from the relationship extraction step using two techniques were evaluated, including extracted entities, relationship types, accuracies of 61.4% with OpenNRE and 87% with REBEL, and processing time. Further, the data processing pipeline was applied to analyze the unstructured dataset from the Transportation Safety Board’s (TSB) Findings for aviation safety analysis. The results showed that structured relationships identified through the pipeline provided valuable indicators, as they captured critical aviation safety information, such as the flight, aircraft type, event, etc. This pipeline can be fine-tuned with a domain-specific knowledge base to provide higher accuracy and better entity detection.

  • Dissertation
  • Cite Count Icon 2
  • 10.33540/1739
Open Information Extraction for Knowledge Representation
  • May 15, 2023
  • Ingy Sarhan

The field of Natural Language Processing (NLP) focuses on developing computational techniques to analyze and extract information from human language. With the exponential growth of unstructured textual data, NLP-based techniques have become essential for extracting valuable insights from this data. However, existing information extraction systems have limitations in terms of extracting valuable information without predefined relations or ontology and storing the extracted knowledge effectively. This Ph.D. thesis aims to enhance open information extraction methods to represent unstructured textual data efficiently and effectively. The first part of the research focuses on Open Information Extraction (OIE) systems and their challenges. Existing OIE methods, including pattern-based and machine learning-based approaches, as well as neural techniques, are analyzed to understand their limitations. A Bidirectional Gated Recurrent Unit (Bi-GRU) OIE model is proposed in Chapter 3, which utilizes contextualized word embeddings to extract relevant triples from unstructured text. Experimental results demonstrate the effectiveness of this model in generating high-quality relation triples. Chapter 4 addresses the lack of labeled data, a common problem in NLP tasks. The research extends the OIE model from Chapter 3 by using learned features to generate relation triples and explores the transferability of these features across different OIE domains and the related task of Relation Extraction (RE). The results show comparable performance with traditional training, indicating the potential of OIE in achieving NLP performance without labeled data. In Chapter 5, the focus shifts to enhancing pre-trained language models for taxonomy classification. Pre-trained language models often struggle with unseen patterns during inference, and the limited size of annotated data poses a challenge. A two-stage fine-tuning procedure, incorporating data augmentation techniques, is proposed to improve the generalizability of pre-trained models. Experimental results demonstrate strong generalizability on unseen data, with an F1 score of 91.25%. Chapter 6 explores the use of OIE for constructing a knowledge graph, specifically in the context of cyber threat intelligence. Open-CyKG, an open cyber threat intelligence knowledge graph framework, is designed using an attention-based neural OIE model and a Named Entity Recognition (NER) model. Refinement and canonicalization techniques are employed to overcome ambiguity and data redundancy during knowledge graph construction. The results show that querying the constructed knowledge graph can be done efficiently, highlighting the support of OIE in knowledge graph development. The proposed components achieve beyond-state-of-the-art results in terms of OIE performance, NER performance, and knowledge graph canonicalization. The research presented in the previous chapters demonstrates significant improvements in the efficiency and effectiveness of open information extraction methods for representing unstructured textual data. These advancements leverage techniques such as data augmentation, multi-stage fine-tuning, and pre-trained language models. The construction of knowledge graphs, enabled by OIE, has the potential to mimic human intelligence and benefit various complex applications, including recommender systems, search engines, and dialog systems.

  • Research Article
  • Cite Count Icon 2
  • 10.1360/ssi-2021-0217
Knowledge graph completion based on parsing graph embedding and a weighted graph convolutional network
  • Nov 1, 2022
  • SCIENTIA SINICA Informationis
  • 妹秋 罗 + 5 more

Knowledge graph completion is an important research issue in knowledge graph construction, knowledge engineering, and natural language processing. A knowledge graph is a knowledge support for realizing accurate knowledge services in general and professional fields. It is also an important breakthrough foundation in information retrieval, question-and-answer interactions, and information recommendation. The low quality and small scale of the knowledge graph are the main bottlenecks that hinder its wide applications. The purpose of knowledge graph completion is to build a large-scale and high-quality knowledge graph for continuously updating and expanding the knowledge graph. Aiming at the difficulty of knowledge graph completion methods to extract deep semantic features from auxiliary information, such as unstructured texts, this study proposes a knowledge graph completion method based on parsing graph embedding and weighted graph convolutional network. This method uses the weighted graph convolutional network to model the semantic dependency parsing of the entity description and construct the semantic dependency parsing graph embedding. Furthermore, it introduces a multi-grained sentence-embedding generation method of the entity description, which is intended to build entity representation learning that can capture multi-grained semantics and deep-level semantic features. The experimental results on two public datasets show that the proposed knowledge graph completion approach outperforms the existing methods, thereby demonstrating its effectiveness and superiority.

  • Conference Article
  • Cite Count Icon 34
  • 10.1145/3300115.3309531
Knowledge Graph based Learning Guidance for Cybersecurity Hands-on Labs
  • May 9, 2019
  • Yuli Deng + 4 more

Hands-on practice is a critical component of cybersecurity education. Most of the existing hands-on exercises or labs materials are usually managed in a problem-centric fashion, while it lacks a coherent way to manage existing labs and provide productive lab exercising plans for cybersecurity learners. With the advantages of big data and natural language processing (NLP) technologies, constructing a large knowledge graph and mining concepts from unstructured text becomes possible, which motivated us to construct a machine learning based lab exercising plan for cybersecurity education. In the research presented by this paper, we have constructed a knowledge graph in the cybersecurity domain using NLP technologies including machine learning based word embedding and hyperlink-based concept mining. We then utilized the knowledge graph during the regular learning process based on the following approaches: 1. We constructed a web-based front-end to visualize the knowledge graph, which allows students to browse and search cybersecurity-related concepts and the corresponding interdependence relations; 2. We created a personalized knowledge graph for each student based on their learning progress and status; 3. We built a personalized lab recommendation system by suggesting more relevant labs based on students' past learning history to maximize their learning outcomes. To measure the effectiveness of the proposed solution, we have conducted a use case study and collected survey data from a graduate-level cybersecurity class. Our study shows that, by leveraging the knowledge graph for the cybersecurity area study, students tend to benefit more and show more interests in cybersecurity area.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 48
  • 10.3390/info13110526
Building Knowledge Graphs from Unstructured Texts: Applications and Impact Analyses in Cybersecurity Education
  • Nov 4, 2022
  • Information
  • Garima Agrawal + 4 more

Knowledge graphs gained popularity in recent years and have been useful for concept visualization and contextual information retrieval in various applications. However, constructing a knowledge graph by scraping long and complex unstructured texts for a new domain in the absence of a well-defined ontology or an existing labeled entity-relation dataset is difficult. Domains such as cybersecurity education can harness knowledge graphs to create a student-focused interactive and learning environment to teach cybersecurity. Learning cybersecurity involves gaining the knowledge of different attack and defense techniques, system setup and solving multi-facet complex real-world challenges that demand adaptive learning strategies and cognitive engagement. However, there are no standard datasets for the cybersecurity education domain. In this research work, we present a bottom-up approach to curate entity-relation pairs and construct knowledge graphs and question-answering models for cybersecurity education. To evaluate the impact of our new learning paradigm, we conducted surveys and interviews with students after each project to find the usefulness of bot and the knowledge graphs. Our results show that students found these tools informative for learning the core concepts and they used knowledge graphs as a visual reference to cross check the progress that helped them complete the project tasks.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.1155/2022/3616432
Correlation between the Dissemination of Classic English Literary Works and Cultural Cognition in the New Media Era
  • Jul 20, 2022
  • Advances in Multimedia
  • Weiwei Guo

With the continuous development of new media technology, the spiritual needs of the masses have been greatly satisfied and the aesthetic ability has also been significantly improved compared with the past. From the current point of view, “literary works,” as the spiritual food of contemporary people, are promoting social spirit. The use of natural language processing and knowledge graph technology can improve cultural cognition to promote the dissemination and development of classic English literature, which has become a necessary means of dissemination of classic English literature. Most of the existing classic English literary works are appreciated based on modern literature datasets. Nowadays, with the continuous development of new media technology, there are fewer studies on the dissemination and cultural cognition of classic English literary works. This makes it impossible for readers to obtain cultural cognition from classic English literary works, making it difficult for the dissemination and development of classic English literary works. In view of the above problems, using natural language processing and knowledge graph technology, taking Shakespeare's play “Hamlet” represented by classic English literary works as an example, the research on the construction method of knowledge graph is carried out and the cultural characteristics in literary works are extracted and analyzed. In parsing, a bidirectional gated recurrent unit network model based on hybrid character embedding is proposed. Based on n-gram embedding, by combining pretraining embedding and radical embedding, it can fully consider the rich semantic information in English literature works to extract. Feature: in terms of named entity recognition, based on the existing iterative atrous convolutional network model, an iterative atrous convolutional network model is proposed. To get the best sequence label and get the last labeled entity information, in terms of knowledge graph construction and visual query, a workflow method for building knowledge graph from unstructured text is proposed and a flask-based knowledge graph visual query system is designed, which applies the best model of the above two tasks. We decode the complete “Hamlet” text, extract entities and their semantic links as nodes and relationships in the knowledge graph, store knowledge through the graph database, and finally form a visual query system that combines the front and back end.

  • Conference Article
  • 10.1109/icbk50248.2020.00062
EntityLDA: A Topic Model for Entity Retrieval on Knowledge Graph
  • Aug 1, 2020
  • Yu Hong + 2 more

Encoding great aof information, knowledge graph (KG) has become a popular data source for information retrieval, especially the entity retrieval task. However, many online KGs include both structured triples and unstructured texts, which makes it difficult to represent entities in a unified form. Moreover, there is also a vocabulary gap between queries given by users and triples contained in KG. To solve these problems, we propose EntityLDA, a topic model which jointly models structured and unstructured parts of KG in order to get complete descriptions of entities. It also bridges the vocabulary gap between users and KG by connecting related words with shared topics. We further propose a retrieval solution based on EntityLDA to retrieve entities under different circumstances. Experimental results show that EntityLDA outperforms baselines in both quantity and quality.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant