Abstract
Climate observations and model simulations produce vast amounts of data. The unprecedented data volume and the complexity of geospatial statistics and analysis requires efficient analysis of big climate data to investigate global problems such as climate change, natural disasters, diseases, and other environmental issues. This paper introduces a high performance query analytical framework to tackle these challenges by leveraging Hive and cloud computing technologies. With this framework, we propose grid transformation, a new perspective for complex climate analysis that applies a series of atomic transformations to terabytes of climate data using SQL-style query (HiveQL). Specifically, we introduce four types of grid transformations (temporal, spatial, local, and arithmetic) to support a broad range of climate analyses, from the basic spatiotemporal aggregation to more sophisticated anomaly detection. Each query is processed as MapReduce tasks in a highly scalable Hadoop cluster as the parallel processing engine. Big climate data are directly stored and managed in a Hadoop Distributed File System without any data format conversion. A prototype is developed to evaluate the feasibility and performance of the framework. Experimental results show that complex and data-intensive climate analysis can be conducted using intuitive SQL queries with good flexibility and performance. This research provides a building block and practical insights in establishing a cyberinfrastructure that provides a high performance and collaborative environment for data-intensive geospatial applications in climate science.
Talk to us
Join us for a 30 min session where you can share your feedback and ask us any queries you have
Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.