Abstract

One challenge of studying cloud workload traces is the lack of available users’ identities. Therefore, clustering methods were used to address this challenge through extracting these identities from workload traces. For better extraction, it is beneficial to select attributes (columns in the traces) for clustering by using feature selection methods. However, the use of general selection methods requires details that are not available for workload traces (e.g. predefined number of clusters). Therefore, in this paper, we present an unsupervised feature selection method for cloud workload traces to identify good candidate attributes for clustering. This method uses Silhouette coefficients to rank attributes that are best for users’ extraction through clustering. The performance of our SeQual method is evaluated in comparison with commonly used (supervised and unsupervised) feature selection methods with the help of clustering quality metrics (i.e. adjusted rand index, entropy and precision). The results show that the SeQual method can compete with the supervised methods and perform better than unsupervised ones, with an average accuracy between 90% and 99%.

Full Text
Paper version not known

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.