Shallow Parsing Research Articles

Commonsense knowledge acquisition and reasoning have long been a core artificial intelligence problem. However, in the past, there has been a lack of scalable methods to collect commonsense knowledge. In this paper, we propose to develop principles for collecting commonsense knowledge based on selectional preference, which is a common phenomenon in human languages that has been shown to be related to semantics. We generalize the definition of selectional preference from one-hop linguistic syntactic relations to higher-order relations over linguistic graphs. Unlike previous commonsense knowledge definitions (e.g., ConceptNet), the selectional preference (SP) knowledge only relies on statistical distributions over linguistic graphs, which can be efficiently and accurately acquired from the unlabeled corpora with modern tools, rather than human-defined relations. As a result, acquiring SP knowledge is a much more scalable way of acquiring commonsense knowledge. Following this principle, we develop a large-scale eventuality (a linguistic term covering activity, state, and event)-based knowledge graph ASER, where each eventuality is represented as a dependency graph, and the relation between them is a discourse relation defined in shallow discourse parsing. The higher-order selectional preference over collected linguistic graphs reflects various kinds of commonsense knowledge. For example, dogs are more likely to bark than cats as the eventuality “dog barks” appears 14,998 times in ASER while “cat barks” only appears 6 times. “Be hungry” is more likely to be the reason rather than result of “eat food” as the edge 〈“be hungry,” Cause, “eat food”〉 appears in ASER while 〈“eat food,” Cause, “be hungry”〉 does not. Moreover, motivated by the observation that humans understand events by abstracting the observed events to a higher level and can thus transfer their knowledge to new events, we propose a conceptualization module on top of the collected knowledge to significantly boost the coverage of ASER. In total, ASER contains 648 million edges between 438 million eventualities. After conceptualization with Probase, a selectional preference based concept-instance relational knowledge base, our concept graph contains 15 million conceptualized eventualities and 224 million edges between them. Detailed analysis is provided to demonstrate its quality. All the collected data, APIs, and tools that can help convert collected SP knowledge into the format of ConceptNet are available at https://github.com/HKUST-KnowComp/ASER.

Read full abstract

Electronic medical records (EMR) hold potential for transformative improvements in the quality and efficiency of healthcare delivery. However, the unstructured nature of EMR data often necessitates manual review and is a significant barrier to leveraging it for downstream analysis. Methods to address this challenge include natural language processing (NLP) algorithms such as tokenization, shallow parsing, and boundary/negation detection. Recent radiation oncology (RO) specific attempts to use NLP to structure EMR data have involved identifying common toxicity terms in on-treatment visit (OTV) notes. Classical NLP works well for positive toxicity term identification (i.e., toxicity present) but is suboptimal for negated symptoms (i.e., toxicity absent). Convoluted neural networks (CNN) aid context detection, however, no publicly available RO-specific CNNs which identify toxicity terms exist. We hypothesized that RO-specific CNNs, naively trained on OTV notes, could improve tabulation of positive and negative toxicity information to an accuracy level suitable for automation.OTV notes (n = 3789) for prostate cancer (PCa) patients treated at our institution between 2019-2021 were identified and analyzed for inclusion/exclusion/omission of CTCAE toxicity terms. The top 5 terms identified (fatigue, nausea, diarrhea, dysuria, hematuria) were used for further analysis. Each note was manually classified by explicit positive identification, negation, or omission of each toxicity term, and used to train an in-house, toxicity-term-specific CNNs. Algorithms were structured as 3-group multiclass classification problem to distinguish between positive or negative symptom identification and omission of the term. Gold standard accuracy measurements were determined using manual review scores of OTV notes for the presence, absence or omission of toxicity terms in each OTV note. Overall, out of sample accuracy and F1 score were determined using a test/train split of OTV notes.The 3-class accuracy of CNNs for the top 5 CTCAE terms present/absent/negated in PCa OTV notes were: fatigue (accuracy = 0.93, F1 = 0.95), diarrhea (accuracy = 0.95, F1 = 0.94), nausea (accuracy = 0.98, F1 = 1.0), dysuria (accuracy = 0.97, F1 = 0.97), hematuria (accuracy = 0.99, F1 = 0.96).Training naïve CNNs with RO-specific training data from OTV notes increased the accuracy of CTCAE toxicity coding. This approach addresses challenges previously encountered using classical NLP from RO EMR data. Therefore, use of CNNs in NLP may reduce barriers to implementation of automated methods to improve data extraction for retrospective and prospective analyses.

Read full abstract

Shallow Parsing Research Articles

Related Topics

Articles published on Shallow Parsing

Multi Task Learning Based Shallow Parsing for Indian Languages

Identification of monolingual and code-switch information from English-Kannada code-switch data

Conserving Semantic Unit Information and Simplifying Syntactic Constituents to Improve Implicit Discourse Relation Recognition.

Deep Learning-based Sequence Labeling Tools for Nepali

Toward a shallow discourse parser for Turkish

Indonesian news classification application with named entity recognition approach

Telugu Dependency Treebank

Automatic Generation of Literary Sentences in French

Multilabel Emotion Tagging for Domain-Specific Texts

ASER: Towards large-scale commonsense knowledge acquisition via higher-order selectional preference over eventualities

METHOD OF DOMAIN ONTOLOGY AUTOMATED REPLENISHMENT FOR THE SUPPORT OF NEW TECHNICAL SOLUTIONS SYNTHESIS. PART I

Convolutional Neural Networks Naively Trained on Radiation Oncology-Specific DATA Outperforms Classical Natural Language Processing Approaches for Automated Identification of Common Toxicity Terms

Opinion Target Extraction Using a Shallow Semantic Parsing Framework

Individual Chunking Ability Predicts Efficient or Shallow L2 Processing: Eye-Tracking Evidence From Multiword Units in Relative Clauses.

Learning Context-Aware Convolutional Filters for Implicit Discourse Relation Classification

New approach to the chunk recoginition in Polish

A novel approach for automatic Bengali question answering system using semantic similarity analysis

Representation learning in discourse parsing: A survey

A Comparative Study of Deep and Shallow Parsing Approaches to Automated Grammaticality Evaluation

Rectifying Incorrectly Part of Speech-Tagged Polysemy Words in Kannada Language for Machine Translation

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Shallow Parsing Research Articles

Related Topics

Articles published on Shallow Parsing

Multi Task Learning Based Shallow Parsing for Indian Languages

Identification of monolingual and code-switch information from English-Kannada code-switch data

Conserving Semantic Unit Information and Simplifying Syntactic Constituents to Improve Implicit Discourse Relation Recognition.

Deep Learning-based Sequence Labeling Tools for Nepali

Toward a shallow discourse parser for Turkish

Indonesian news classification application with named entity recognition approach

Telugu Dependency Treebank

Automatic Generation of Literary Sentences in French

Multilabel Emotion Tagging for Domain-Specific Texts

ASER: Towards large-scale commonsense knowledge acquisition via higher-order selectional preference over eventualities

METHOD OF DOMAIN ONTOLOGY AUTOMATED REPLENISHMENT FOR THE SUPPORT OF NEW TECHNICAL SOLUTIONS SYNTHESIS. PART I

Convolutional Neural Networks Naively Trained on Radiation Oncology-Specific DATA Outperforms Classical Natural Language Processing Approaches for Automated Identification of Common Toxicity Terms

Opinion Target Extraction Using a Shallow Semantic Parsing Framework

Individual Chunking Ability Predicts Efficient or Shallow L2 Processing: Eye-Tracking Evidence From Multiword Units in Relative Clauses.

Learning Context-Aware Convolutional Filters for Implicit Discourse Relation Classification

New approach to the chunk recoginition in Polish

A novel approach for automatic Bengali question answering system using semantic similarity analysis

Representation learning in discourse parsing: A survey

A Comparative Study of Deep and Shallow Parsing Approaches to Automated Grammaticality Evaluation

Rectifying Incorrectly Part of Speech-Tagged Polysemy Words in Kannada Language for Machine Translation