An approach for text extraction from web news page

Hu Mingsheng,Jia Zhijuan,Zhang Xiangyu

doi:10.1109/isra.2012.6219250

Abstract

With the rapid development of Internet and Web technology, Web page has become a main carrier of information publishing. In connection with the problems of current complex implementation, high error rate and low extraction speed of Web information extraction technology, this paper proposes a new method of Web extraction based on the characteristics of structure of Web page. This method is to use tree structure of DOM (Document Object Model) when analyzing web page, parsing the Web page into DOM tree to sort the scattered web pages, by the using of the characteristics of Chinese web pages similar in information structure and aggregated distribution to achieve simply with good versatility. At the same time, this method can reduce the complexity when dealing with the structure of web page and increase the speed of the Web information extraction. At present, the method has been applied to the news page automatic classification system, which is good to meet the system's requirements.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An approach for text extraction from web news page

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

X-BROT: Prototyping of Compatibility Testing Tool for Web Application Based on Document Analysis Technology
Hiroshi Tanaka
-
Hiroshi TanakaHiroshi Tanaka
01 Sep 2019
01 Sep 2019

NLP based intelligent news search engine using information extraction from e-newspapers
Monisha Kanakaraj ... S Sowmya Kamath
-
Monisha Kanakaraj, et. al.Monisha Kanakaraj ... S Sowmya Kamath
01 Dec 2014
01 Dec 2014

Structure Analysis of Tor Hidden Services Using DOM-Inspired Graphs
Ashwini Dalvi ... Kunjal Shah
-
Ashwini Dalvi, et. al.Ashwini Dalvi ... Kunjal Shah
06 May 2022
06 May 2022

Extracting news content with visual unit of web pages
Wenhao Zhu ... Zhiguo Lu
-
Wenhao Zhu, et. al. Wenhao Zhu ... Zhiguo Lu
01 Jun 2015
01 Jun 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An approach for text extraction from web news page

Abstract

Talk to us

Similar Papers