Version Of C4 Research Articles

In this work, we have considered the ensemble of classifier chains (ECC) algorithm in order to solve the multi-label classification (MLC) task. It starts from binary relevance algorithm (BR), a simple and direct approach to MLC that has been shown to provide good results in practice. Nevertheless, unlike BR, ECC aims to exploit the correlations between labels. ECC uses an algorithm of traditional supervised classification in order to approach the binary problems. Within this field, Credal C4.5 (CC4.5) is a new version of the well-known C4.5 algorithm that uses imprecise probabilities in order to estimate the probability distribution of the class variable. This new version of C4.5 algorithm has been shown to provide better performance when noisy datasets are classified. In MLC, the intrinsic noise might be higher than in traditional supervised classification. The reason is very simple: in MLC, there are multiple labels, whereas in traditional classification there is just a class variable. Thus, there is more probability of error for an instance. For the previous reasons, the performance of ECC with CC4.5 as base classifier is studied in this work. We have carried out an extensive experimental analysis with several multi-label datasets, different noise levels and a large number of evaluation metrics for MLC. This experimental study has shown that, generally, ECC has better performance with CC4.5 as base classifier than using C4.5. The higher is the label noise level introduced in the data, the more significative is this improvement. Therefore, it is probably suitable to use imprecise probabilities in Decision Trees within MLC.

Read full abstract

Traditionally, the performance of a classifier is measured by its classification accuracy or error rate. In fact, probability-based classifiers also produce the class probability estimation (the probability that a test instance belongs to the predicted class). This information is often ignored in classification, as long as the class with the highest class probability estimation is identical to the actual class. In many data mining applications, however, classification accuracy and error rate are not enough. For example, in direct marketing, we often need to deploy different promotion strategies to customers with different likelihood (class probability) of buying some products. Thus, accurate class probability estimations are often required to make optimal decisions. In this paper, we firstly review some state-of-the-art probability-based classifiers and empirically investigate their class probability estimation performance. From our experimental results, we can draw a conclusion: C4.4 is an attractive algorithm for class probability estimation. Then, we present a locally weighted version of C4.4 to scale up its class probability estimation performance by combining locally weighted learning with C4.4. We call our improved algorithm locally weighted C4.4, simply LWC4.4. We experimentally test LWC4.4 using the whole 36 UCI data sets selected by Weka. The experimental results show that LWC4.4 significantly outperforms C4.4 in terms of class probability estimation.

Read full abstract

Version Of C4 Research Articles

Articles published on Version Of C4

Using Credal C4.5 for Calibrated Label Ranking in Multi-Label Classification

Ensemble of classifier chains and Credal C4.5 for solving multi-label classification

DECISION TREE WITH BETTER CLASS PROBABILITY ESTIMATION

A genetic-algorithm for discovering small-disjunct rules in data mining

TWO-WAY INDUCTION

Improved Use of Continuous Attributes in C4.5

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Version Of C4 Research Articles

Articles published on Version Of C4

Using Credal C4.5 for Calibrated Label Ranking in Multi-Label Classification

Ensemble of classifier chains and Credal C4.5 for solving multi-label classification

DECISION TREE WITH BETTER CLASS PROBABILITY ESTIMATION

A genetic-algorithm for discovering small-disjunct rules in data mining

TWO-WAY INDUCTION

Improved Use of Continuous Attributes in C4.5