Spark for Social Science

Graham Macdonald,Jeffrey Levy,Alex Engler,Sarah Armstrong

doi:10.23889/ijpds.v3i5.1044

Spark for Social Science

Graham Macdonald, Jeffrey Levy + Show 2 more

Open Access

https://doi.org/10.23889/ijpds.v3i5.1044

Copy DOI

Journal: International Journal of Population Data Science	Publication Date: Oct 10, 2018
License type: CC BY-NC-ND 4.0

Affiliation: Urban Institute, University of Chicago

#Amazon Web Services #Elastic MapReduce + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

Urban has developed an elastic and powerful approach to the analysis of massive datasets using Amazon Web Services’ Elastic MapReduce (EMR) and the Spark framework for distributed memory and processing. The goal of the project is to deliver powerful and elastic Spark clusters to researchers and data analysts with as little setup time and effort possible, and at low cost. To do that, at the Urban Institute, we use two critical components: (1) an Amazon Web Services (AWS) CloudFormation script to launch AWS Elastic MapReduce (EMR) clusters (2) a bootstrap script that runs on the Master node of the new cluster to install statistical programs and development environments (RStudio and Jupyter Notebooks). The Urban Institute’s Spark for Social Science Github page holds code used to setup the cluster and tutorials for learning how to program in R and Python.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: International Journal of Population Data Science

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.