MiSoSouP

Matteo Riondato,Fabio Vandin

doi:10.1145/3385653

Abstract

We present MiSoSouP, a suite of algorithms for extracting high-quality approximations of the most interesting subgroups, according to different popular interestingness measures, from a random sample of a transactional dataset. We describe a new formulation of these measures as functions of averages, that makes it possible to approximate them using sampling. We then discuss how pseudodimension, a key concept from statistical learning theory, relates to the sample size needed to obtain an high-quality approximation of the most interesting subgroups. We prove an upper bound on the pseudodimension of the problem at hand, which depends on characteristic quantities of the dataset and of the language of patterns of interest. This upper bound then leads to small sample sizes. Our evaluation on real datasets shows that MiSoSouP outperforms state-of-the-art algorithms offering the same guarantees, and it vastly speeds up the discovery of subgroups w.r.t. analyzing the whole dataset.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MiSoSouP

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Knowledge Discovery from Data

Lead the way for us

Journal: ACM Transactions on Knowledge Discovery from Data	Publication Date: Jun 21, 2020
Citations: 5

Similar Papers

MiSoSouP
Matteo Riondato ... Fabio Vandin
-
Matteo Riondato, et. al.Matteo Riondato ... Fabio Vandin
19 Jul 2018
19 Jul 2018

TipTap: Approximate Mining of Frequent k -Subgraph Patterns in Evolving Graphs
Muhammad Anis Uddin Nasir ... Gianmarco De Francisci Morales
ACM Transactions on Knowledge Discovery from Data | VOL. 15
Muhammad Anis Uddin Nasir, et. al.Muhammad Anis Uddin Nasir ... Gianmarco De Francisci Morales
21 Apr 2021
ACM Transactions on Knowledge Discovery from Data | VOL. 15

Statistical Foundations of Data Science
Jianqing Fan ... Cun-Hui Zhang
-
Jianqing Fan, et. al.Jianqing Fan ... Cun-Hui Zhang
20 Sep 2020
20 Sep 2020

Unsupervised machine learning for the discovery of latent disease clusters and patient subgroups using electronic health records.
Yanshan Wang ... Ahmad P Tafti
Journal of Biomedical Informatics | VOL. 102
Yanshan Wang, et. al.Yanshan Wang ... Ahmad P Tafti
28 Dec 2019
Journal of Biomedical Informatics | VOL. 102

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MiSoSouP

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Knowledge Discovery from Data