Effective Job Execution in Hadoop Over Authorized Deduplicated Data

Sachin Arun Thanekar,A.B Bagwan,K Subrahmanyam

doi:10.14704/web/v17i2/web17043

Abstract

Existing Hadoop treats every job as an independent job and destroys metadata of preceding jobs. As every job is independent, again and again it has to read data from all Data Nodes. Moreover relationships between specific jobs are also not getting checked. Lack of Specific user identities creation and forming groups, managing user credentials are the weaknesses of HDFS. Due to which overall performance of Hadoop becomes very poor. So there is a need to improve the Hadoop performance by reusing metadata, better space management, better task execution by checking deduplication and securing data with access rights specification. In our proposed system, task deduplication technique is used. It checks the similarity between jobs by checking block ids. Job metadata and data locality details are stored on Name Node which results in better execution of job. Metadata of executed jobs is preserved. Thus by preserving job metadata re computations time can be saved. Experimental results show that there is an improvement in job execution time, reduced storage space. Thus, improves Hadoop performance.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Effective Job Execution in Hadoop Over Authorized Deduplicated Data

Abstract

Talk to us

Similar Papers

More From: Webology

Lead the way for us

Journal: Webology	Publication Date: Dec 21, 2020
License type: cc-by-nc-nd

Similar Papers

Locality-Aware and Energy-Aware Job Pre-Assignment for Mapreduce
Lei Chen ... Tinqing He
-
Lei Chen, et. al.Lei Chen ... Tinqing He
01 Sep 2016
01 Sep 2016

SHadoop: Improving MapReduce performance by optimizing job execution mechanism in Hadoop clusters
Rong Gu ... Yihua Huang
Journal of Parallel and Distributed Computing | VOL. 74
Rong Gu, et. al.Rong Gu ... Yihua Huang
01 Dec 2013
Journal of Parallel and Distributed Computing | VOL. 74

A Comprehensive Framework for Specifying Clairvoyance, Constraints and Periodicity in Real-Time Scheduling
K Subramani
The Computer Journal | VOL. 48
K SubramaniK Subramani
01 Mar 2005
The Computer Journal | VOL. 48

Analyzing & optimizing hadoop performance
Ankita Jain ... Monika Choudhary
-
Ankita Jain, et. al.Ankita Jain ... Monika Choudhary
01 Mar 2017
01 Mar 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Effective Job Execution in Hadoop Over Authorized Deduplicated Data

Abstract

Talk to us

Similar Papers

More From: Webology