THE USE OF ROUGH CLASSIFICATION AND TWO THRESHOLD TWO DIVISORS FOR DEDUPLICATION

Hashem B Jehlol,Loay E George

doi:10.25195/ijci.v49i1.379

Abstract

The data deduplication technique efficiently reduces and removes redundant data in big data storage systems. The main issue is that the data deduplication requires expensive computational effort to remove duplicate data due to the vast size of big data. The paper attempts to reduce the time and computation required for data deduplication stages. The chunking and hashing stage often requires a lot of calculations and time. This paper initially proposes an efficient new method to exploit the parallel processing of deduplication systems with the best performance. The proposed system is designed to use multicore computing efficiently. First, The proposed method removes redundant data by making a rough classification for the input into several classes using the histogram similarity and k-mean algorithm. Next, a new method for calculating the divisor list for each class was introduced to improve the chunking method and increase the data deduplication ratio. Finally, the performance of the proposed method was evaluated using three datasets as test examples. The proposed method proves that data deduplication based on classes and a multicore processor is much faster than a single-core processor. Moreover, the experimental results showed that the proposed method significantly improved the performance of Two Threshold Two Divisors (TTTD) and Basic Sliding Window BSW algorithms.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

THE USE OF ROUGH CLASSIFICATION AND TWO THRESHOLD TWO DIVISORS FOR DEDUPLICATION

Abstract

Talk to us

Similar Papers

More From: Iraqi Journal for Computers and Informatics

Lead the way for us

Journal: Iraqi Journal for Computers and Informatics	Publication Date: Jun 11, 2023
License type: CC BY-NC-ND 4.0

Similar Papers

Detection of duplicated data with minimum overhead and secure data transmission for sensor big data
S Beulah ... F Ramesh Dhanaseelan
Cluster Computing | VOL. 22
S Beulah, et. al.S Beulah ... F Ramesh Dhanaseelan
21 Aug 2017
Cluster Computing | VOL. 22

Privacy-preserving smart IoT-based healthcare big data storage and self-adaptive access control system
Yang Yang ... Victor Chang
Information Sciences | VOL. 479
Yang Yang, et. al.Yang Yang ... Victor Chang
06 Feb 2018
Information Sciences | VOL. 479

On managing geospatial big-data in emergency management
Kuien Liu ... Danhuai Guo
-
Kuien Liu, et. al.Kuien Liu ... Danhuai Guo
03 Nov 2015
03 Nov 2015

Data Deduplication Techniques for Big Data Storage Systems
Niteesha Sharma ... Dr A V Krishna Prasad
International Journal of Innovative Technology and Exploring Engineering | VOL. 8
Niteesha Sharma, et. al.Niteesha Sharma ... Dr A V Krishna Prasad
30 Aug 2019
International Journal of Innovative Technology and Exploring Engineering | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

THE USE OF ROUGH CLASSIFICATION AND TWO THRESHOLD TWO DIVISORS FOR DEDUPLICATION

Abstract

Talk to us

Similar Papers

More From: Iraqi Journal for Computers and Informatics