Abstract

For any educational project, it is important and challenging to know, at the moment of enrollment, whether a given student is likely to successfully pass the academic year. This task is not simple at all because many factors contribute to college failure. Being able to infer how likely is an enrolled student to present promotions problems, is undoubtedly an interesting challenge for the areas of data mining and education. In this paper, we propose the use of data mining techniques in order to predict how likely a student is to succeed in the academic year. Normally, there are more students that success than fail, resulting in an imbalanced data representation. To cope with imbalanced data, we introduce a new algorithm based on probabilistic Rough Set Theory (RST). Two ideas are introduced. The first one is the use of two different threshold values for the similarity between objects when dealing with minority or majority examples. The second idea combines the original distribution of the data with the probabilities predicted by the RST method. Our experimental analysis shows that we obtain better results than a range of state-of-the-art algorithms.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call