Available techniques in hadoop small file issue

M B Masadeh,M S Azmi,S S S Ahmad

doi:10.11591/ijece.v10i2.pp2097-2101

M B Masadeh, M S Azmi + Show 1 more

Open Access

https://doi.org/10.11591/ijece.v10i2.pp2097-2101

Copy DOI

Abstract

Hadoop is an optimal solution for big data processing and storing since being released in the late of 2006, hadoop data processing stands on master-slaves manner [1] that’s splits the large file job into several small files in order to process them separately, this technique was adopted instead of pushing one large file into a costly super machine to insights some useful information. Hadoop runs very good with large file of big data, but when it comes to big data in small files it could facing some problems in performance, processing slow down, data access delay, high latency and up to a completely cluster shutting down [2]. In this paper we will high light on one of hadoop’s limitations, that’s affects the data processing performance, one of these limits called “big data in small files” accrued when a massive number of small files pushed into a hadoop cluster which will rides the cluster to shut down totally. This paper also high light on some native and proposed solutions for big data in small files, how do they work to reduce the negative effects on hadoop cluster, and add extra performance on storing and accessing mechanism.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: International Journal of Electrical and Computer Engineering (IJECE)	Publication Date: Apr 1, 2020
Citations: 2	License type: CC BY-NC 4.0

R Discovery Prime

R Discovery Prime

Available techniques in hadoop small file issue

Abstract

Talk to us

Similar Papers

More From: International Journal of Electrical and Computer Engineering (IJECE)

Lead the way for us

Similar Papers

Survey on Resource Management Solutions to Speed up Processing Small Files in Hadoop Cluster
Prof Shwetha K S ... Dr Chandramouli H
International Journal of Scientific Research in Science, Engineering and Technology | VOL. -
Prof Shwetha K S, et. al. Prof Shwetha K S ... Dr Chandramouli H
05 Nov 2023
International Journal of Scientific Research in Science, Engineering and Technology | VOL. -

<title>Performance evaluation of a high-speed switched network for PACS</title>
Randy H Zhang ... Lu J Huang
-
Randy H Zhang, et. al.Randy H Zhang ... Lu J Huang
13 Jul 1998
13 Jul 1998

Improving Small File Management in Hadoop
Samira Khoulji ... M.L Kerkeb
Transactions on Machine Learning and Artificial Intelligence | VOL. 5
Samira Khoulji, et. al.Samira Khoulji ... M.L Kerkeb
31 Aug 2017
Transactions on Machine Learning and Artificial Intelligence | VOL. 5

An effective merge strategy based hierarchy for improving small file problem on HDFS
Zhipeng Gao ... Yinghao Qin
-
Zhipeng Gao, et. al.Zhipeng Gao ... Yinghao Qin
01 Aug 2016
01 Aug 2016

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Available techniques in hadoop small file issue

Abstract

Talk to us

Similar Papers

More From: International Journal of Electrical and Computer Engineering (IJECE)