Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System

Bok-Keun Sun

doi:10.9708/jksci.2016.21.4.055

Abstract

In this paper, we propose the INDEM(Internet News Data Extraction Middleware) system for the removal of the unnecessary data in internet news. Although data on the internet can be used in various fields such as source of data of IR(Information Retrieval), Data mining and knowledge information service, it contains a lot of unnecessary information. The removal of the unnecessary data is a problem to be solved prior to the study of the knowledge-based information service that is based on the data of the web page. The INDEM system parses html and explores the XPath, and it is to perform the analysis. The user simply utilize INDEM by implementing an abstract class that provides INDEM, and can obtain the analysis information. INDEM System through this process delivers the analysis information including the main contents of news site to the users. In this paper, the INDEM system was adapted in a stand-alone and web service system and it was evaluated on the basis of 16 news site. As a result, performance of the INDEM system is affected in html source data size and complexity of used html grammar than the main news data size.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System

Abstract

Talk to us

Similar Papers

More From: Journal of the Korea Society of Computer and Information

Lead the way for us

Similar Papers

Knowledge Elicitation through Web-Based Data Mining Services
Shonali Krishnaswamy ... Arkady Zaslavsky
-
Shonali Krishnaswamy, et. al.Shonali Krishnaswamy ... Arkady Zaslavsky
01 Jan 2001
01 Jan 2001

Data-Aware Web Service Recommender System for Energy-Efficient Data Mining Services
Zainab Al-Zanbouri ... Chen Ding
-
Zainab Al-Zanbouri, et. al.Zainab Al-Zanbouri ... Chen Ding
01 Nov 2018
01 Nov 2018

The Analysis of Internet Commercial Judicial Based on Big Data Alliance and Mining Service Process Model
Zhao Zhonglong ... Wang Hongliang
Complexity | VOL. 2021
Zhao Zhonglong, et. al.Zhao Zhonglong ... Wang Hongliang
01 Jan 2020
Complexity | VOL. 2021

A Review on Big Data Mining in Cloud Computing
Bhaludra R Nadh Singh ... B Raja Srinivasa Reddy
-
Bhaludra R Nadh Singh, et. al.Bhaludra R Nadh Singh ... B Raja Srinivasa Reddy
01 Jan 2017
01 Jan 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System

Abstract

Talk to us

Similar Papers

More From: Journal of the Korea Society of Computer and Information