XBA1-GB 8346
Big Data
New York University · UGRD · Fall 2026
Catalog description
This course offers an in-depth hands-on exploration of cutting-edge cloud technologies used for big data analytics. The pre-module will cover background readings on the theoretical foundations of Hadoop and MapReduce, as well as business articles on how Hadoop and related technologies are used by companies. The pre-module will also cover some basics of navigating Google Cloud Platform (GCP) for uploading and analyzing data using Google Cloud Storage, BigQuery, and PySpark on Dataproc (Hadoop cluster). Next, in the in-class module, we will focus on the Hadoop Big Data environment – specifically, Linux, Hadoop distributed file system (HDFS), Apache Sqoop, Apache Pig and Apache Hive – for data management and extract-transform-load (ETL) operations. Then, we will use Google Cloud Platform (GCP) to do hands-on exercises to upload and process large files (in the 1 GB+ range) for your Capstone project and other purposes. On GCP we will focus on cloud file storage, querying cloud data, visualizing cloud data, and using PySpark to run analytics on cloud data. In the post-module, we will have assignments on the hands-on material covered in the in-class module.
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verified- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff