Multiple clusterings: Recent advances and perspectives

Guoxian Yu,Liangrui Ren,Jun Wang,Carlotta Domeniconi,Xiangliang Zhang

doi:10.1016/j.cosrev.2024.100621

Abstract

Clustering is a fundamental data exploration technique to discover hidden grouping structure of data. With the proliferation of big data, and the increase of volume and variety, the complexity of data multiplicity is increasing as well. Traditional clustering methods can provide only a single clustering result, which restricts data exploration to one single possible partition. In contrast, multiple clustering can simultaneously or sequentially uncover multiple non-redundant and distinct clustering solutions, which can reveal multiple interesting hidden structures of the data from different perspectives. For these reasons, multiple clustering has become a popular and promising field of study. In this survey, we have conducted a systematic review of the existing multiple clustering methods. Specifically, we categorize existing approaches according to four different perspectives (i.e., multiple clustering in the original space, in subspaces and on multi-view data, and multiple co-clustering). We summarize the key ideas underlying the techniques and their objective functions, and discuss the advantages and disadvantages of each. In addition, we built a repository of multiple clustering resources (i.e., benchmark datasets and codes). Finally, we discuss the key open issues for future investigation.

Full Text