Erasable Virtual HyperLogLog for Approximating Cumulative Distribution over Data Streams

Peng Jia,Junzhou Zhao,Ye Yuan,Pinghui Wang,Jing Tao,Xiaohong Guan

doi:10.1109/tkde.2021.3052938

Abstract

Many real-world datasets are given in the stream of entity-identifier pairs, and measuring data distribution on these datasets is fundamental for applications such as privacy protection. In this paper, we study the problem of computing the cumulative distribution for different cardinalities (i.e., the number of distinct entities owning the same identifier). However, previous sketch-based methods cost large memory space especially when there are a large number of identifiers, and sampling-based methods require much time for cardinality estimation. A recent work KHyperLogLog combines both sketch and sampling methods but it is wasteful to separately build a HyperLogLog sketch of large size for identifiers with small cardinalities. To address these challenges, we propose a memory-efficient method EV-HLL, which designs a shared structure to store all sampled identifiers and their entities and utilizes additional sketches to track value updates during the sampling procedure. Meanwhile, EV-HLL provides real-time unbiased estimations according to value changes whenever a new entity-identifier pair arrives. We evaluate the performance of EV-HLL and other state-of-the-arts on real-world available datasets. Experimental results demonstrate that comparing to other methods, EV-HLL effectively reduces their memory usage with the same estimation accuracy and has higher accuracy with the same memory usage.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Erasable Virtual HyperLogLog for Approximating Cumulative Distribution over Data Streams

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Knowledge and Data Engineering

Lead the way for us

Journal: IEEE Transactions on Knowledge and Data Engineering	Publication Date: Nov 1, 2022
Citations: 3

Similar Papers

LogLog Filter: Filtering Cold Items within a Large Range over High Speed Data Streams
Peng Jia ... Jing Tao
-
Peng Jia, et. al.Peng Jia ... Jing Tao
01 Apr 2021
01 Apr 2021

Differential Privacy for Protecting Private Patterns in Data Streams
He Gu ... Vera Goebel
-
He Gu, et. al.He Gu ... Vera Goebel
01 Apr 2023
01 Apr 2023

RETRACTED ARTICLE: Comprehensive analysis for class imbalance data with concept drift using ensemble based classification
S Priya ... R Annie Uthra
Journal of Ambient Intelligence and Humanized Computing | VOL. 12
S Priya, et. al.S Priya ... R Annie Uthra
11 Apr 2020
Journal of Ambient Intelligence and Humanized Computing | VOL. 12

A Framework for Classification in Data Streams Using Multi-strategy Learning
Ali Pesaranghader ... Herna L Viktor
-
Ali Pesaranghader, et. al.Ali Pesaranghader ... Herna L Viktor
01 Jan 2015
01 Jan 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Erasable Virtual HyperLogLog for Approximating Cumulative Distribution over Data Streams

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Knowledge and Data Engineering