Clustering news articles using efficient similarity measure and N-grams

Desmond Bala Bisandu,Musa Muhammad Liman,Rajesh Prasad

doi:10.1504/ijkedm.2018.095525

Abstract

The rapid progress of information technology and web makes it easier to store huge amount of collected textual information, e.g., blogs, news articles, e-mail messages, reviews and forum postings. The growing size of textual dataset with high-dimensions and natural language pose a big challenge making it hard for such information to be categorised efficiently. Document clustering is an automatic unsupervised machine learning technique that aimed at grouping related set of items into clusters or subsets. The target is creating clusters with high internal coherence, but different from each other substantially. This paper presents a new document clustering technique using N-grams and efficient similarity measure known as 'improved sqrt-cosine similarity measure'. Comprehensive experiments are conducted to evaluate our proposed clustering technique and compared with an existing method. The results of the experiments show that our proposed clustering technique outperforms the existing techniques.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Clustering news articles using efficient similarity measure and N-grams

Abstract

Talk to us

Similar Papers

More From: International Journal of Knowledge Engineering and Data Mining

Lead the way for us

Journal: International Journal of Knowledge Engineering and Data Mining	Publication Date: Jan 1, 2018
Citations: 21

Similar Papers

Clustering News Articles using Efficient Similarity Measure and N-grams
Rajesh Prasad ... Desmond Bisandu
International Journal of Knowledge Engineering and Data Mining | VOL. 5
Rajesh Prasad, et. al.Rajesh Prasad ... Desmond Bisandu
01 Jan 2018
International Journal of Knowledge Engineering and Data Mining | VOL. 5

A similarity assessment technique for effective grouping of documents
Tanmay Basu ... C.A Murthy
Information Sciences | VOL. 311
Tanmay Basu, et. al.Tanmay Basu ... C.A Murthy
21 Mar 2015
Information Sciences | VOL. 311

Automatic Scientific Document Clustering Using Self-organized Multi-objective Differential Evolution
Naveen Saini ... Pushpak Bhattacharyya
Cognitive Computation | VOL. 11
Naveen Saini, et. al.Naveen Saini ... Pushpak Bhattacharyya
19 Dec 2018
Cognitive Computation | VOL. 11

Development of Document Clustering Technique for Gurmukhi Script using Fuzzy Term Weight
Mukesh Kumar ... Amandeep Verma
International Journal of Recent Technology and Engineering (IJRTE) | VOL. 8
Mukesh Kumar, et. al.Mukesh Kumar ... Amandeep Verma
30 Jul 2019
International Journal of Recent Technology and Engineering (IJRTE) | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Clustering news articles using efficient similarity measure and N-grams

Abstract

Talk to us

Similar Papers

More From: International Journal of Knowledge Engineering and Data Mining