Web Page Classification System Research Articles

Nowadays, the Internet contain s a wide variety of online documents, making finding useful information about a given subject impossible, as well as retrieving irrelevant pages. Web document and page recognition software is useful in a variety of fields, including news, medicine, and fitness, research, and information technology. To enhance search capability, a large number of web page classification methods have been proposed, especially for news web pages. Furthermore existing classification approaches seek to distinguish news web pages while still reducing the high dimensionality of features derived from these pages. Due to the lack of automated classification methods, this paper focuses on the classification of news web pages based on their scarcity and importance. This work will establish different models for the identification and classification of the web pages. The data sets used in this paper were collected from popular news websites. In the research work we have used BBC dataset that has five predefined categories. Initially the input source can be preprocessed and the errors can be eliminated. Then the features can be extracted depend upon the web page reviews using Term frequency-inverse document frequency vectorization. In the work 2225 documents are represented with the 15286 features, which represents the tf-idf score for different unigrams and bigrams. This type of the representation is not only used for classification task also helpful to analyze the dataset. Feature selection is done by using the chi-squared test which will be in the task of finding the terms that are most correlated with each of the categories. Then the pointed features can be selected using chi-squared test. Finally depend upon the classifier the web page can be classified. The results showed that list has obtained the highest percentage, which reflect its effectiveness on the classification of web pages.

Read full abstract

본 논문은 온톨로지(ontology)에 기반 한 자동화된 웹 페이지 분류 시스템을 제안한다. 웹 페이지의 분류를 위하여 첫 번째 단계에서는 각 웹 페이지가 속한 범주(category)를 대표할 수 있는 단어를 선정하며, 이를 위하여 단어빈도와 문서빈도를 곱한 값을 계산한다. 두 번째 단계에서는 첫 번째 단계에 의해 선택된 단어의 정보이득(information gain)을 계산해 분류 확률이 높은 단어를 우선적으로 선정한다. 두 단계를 통하여 선정된 단어들과 웹 페이지의 분류 정보를 가지고, 기계학습에 의하여 컴파일 된 규칙(compiled rules)을 생성한다. 생성된 규칙은 임의의 웹 페이지들을 도메인 온톨로지에 의해 정의된 범주 별로 분류할 수 있도록 한다. 본 논문의 실험에서는 주어진 웹 페이지 집합에서 각 범주 별로 평균 240개의 단어로부터 78개의 단어를 결과적으로 선정하였으며, 이를 바탕으로 웹 페이지 분류 규칙을 생성하였다. 실험 결과에서 제안한 시스템의 평균 분류 정확도는 약 83.52%로 측정되었다. In this paper, we present an automated Web page classification system based upon ontology. As a first step, to identify the representative terms given a set of classes, we compute the product of term frequency and document frequency. Secondly, the information gain of each term prioritizes it based on the possibility of classification. We compile a pair of the terms selected and a web page classification into rules using machine learning algorithms. The compiled rules classify any Web page into categories defined on a domain ontology. In the experiments, 78 terms out of 240 terms were identified as representative features given a set of Web pages. The resulting accuracy of the classification was, on the average, 83.52%.

Read full abstract

Web Page Classification System Research Articles

Related Topics

Articles published on Web Page Classification System

Semantic-based Web Page Classification System Using Enhanced C4.5

Automatic Web Page Classification System with Improved Accuracy

Performance Comparison between Keyword-based and Semantic-based Web Page Classification Systems

Web Page Classification Using RNN

A Naive Bayes approach for URL classification with supervised feature selection and rejection framework

Web Page Advertisement Classification

An ant colony optimization based feature selection for web page classification.

메타 태그를 이용한 자동 웹페이지 분류 시스템

A Web page classification system based on a genetic algorithm using tagged-terms as features

온톨로지 기반의 웹 페이지 분류 시스템

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Web Page Classification System Research Articles

Related Topics

Articles published on Web Page Classification System

Semantic-based Web Page Classification System Using Enhanced C4.5

Automatic Web Page Classification System with Improved Accuracy

Performance Comparison between Keyword-based and Semantic-based Web Page Classification Systems

Web Page Classification Using RNN

A Naive Bayes approach for URL classification with supervised feature selection and rejection framework

Web Page Advertisement Classification

An ant colony optimization based feature selection for web page classification.

메타 태그를 이용한 자동 웹페이지 분류 시스템

A Web page classification system based on a genetic algorithm using tagged-terms as features

온톨로지 기반의 웹 페이지 분류 시스템