Evaluation of novelty metrics for sentence-level novelty mining

Flora S Tsai,Wenyin Tang,Kap Luk Chan

doi:10.1016/j.ins.2010.02.020

Abstract

This work addresses the problem of detecting novel sentences from an incoming stream of text data, by studying the performance of different novelty metrics, and proposing a mixed metric that is able to adapt to different performance requirements. Existing novelty metrics can be divided into two types, symmetric and asymmetric, based on whether the ordering of sentences is taken into account. After a comparative study of several different novelty metrics, we observe complementary behavior in the two types of metrics. This finding motivates a new framework of novelty measurement, i.e. the mixture of both symmetric and asymmetric metrics. This new framework of novelty measurement performs superiorly under different performance requirements varying from high-precision to high-recall as well as for data with different percentages of novel sentences. Because it does not require any prior information, the new metric is very suitable for real-time knowledge base applications such as novelty mining systems where no training data is available beforehand.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Evaluation of novelty metrics for sentence-level novelty mining

Abstract

Talk to us

Similar Papers

More From: Information Sciences

Lead the way for us

Journal: Information Sciences	Publication Date: Feb 24, 2010
Citations: 38

Similar Papers

A stream-sensitive distributed approach for configuring cascaded classifier topologies in real-time large-scale stream mining systems
Abtin Shahkarami ... Nader Bagherzadeh
SN Applied Sciences | VOL. 1
Abtin Shahkarami, et. al.Abtin Shahkarami ... Nader Bagherzadeh
18 May 2019
SN Applied Sciences | VOL. 1

Online supervised hashing
Fatih Cakir ... Stan Sclaroff
-
Fatih Cakir, et. al.Fatih Cakir ... Stan Sclaroff
01 Sep 2015
01 Sep 2015

EMM-CLODS: An Effective Microcluster and Minimal Pruning CLustering-Based Technique for Detecting Outliers in Data Streams
Mohamed Jaward Bah ... Ji Zhang
Complexity | VOL. 2021
Mohamed Jaward Bah, et. al.Mohamed Jaward Bah ... Ji Zhang
13 Sep 2021
Complexity | VOL. 2021

A linear regression-based frequent itemset forecast algorithm for stream data
S Shukla ... B Verma
-
S Shukla, et. al.S Shukla ... B Verma
01 Dec 2009
01 Dec 2009

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Evaluation of novelty metrics for sentence-level novelty mining

Abstract

Talk to us

Similar Papers

More From: Information Sciences