PQHS 486
Data Science Infrastructure for Biomedical Computing
Case Western Reserve University · UGRD · Fall 2026
Catalog description
This advanced course introduces concepts and techniques necessary to construct and configure distributed computing clusters for large-scale biomedical data processing and apply these tools for data analysis. Specifically, this course will cover 1) Linux and network configurations for cluster computing, 2) deployment and configuration of Hadoop, HDFS, and SPARK, and 3) performance assessment and analysis of typical large-scale biomedical data sets. Skills will be taught in a highly interactive workshop style with minimal lectures. During the first two modules of the course, students will be assigned tasks to be completed using command line operations to configure their own Raspberry Pi-style device. After basic configuration, students will transition to group work, combining individual devices into computing clusters. During the last module of the course, students will be assigned typical biomedical data science datasets and will implement solutions for data analysis on the computing cluster. Concepts learned will be directly relevant and translatable to clusters running in cloud computing environments. This course will follow a modified IQ-style format which requires both independent and group learning. Course topics will be distributed among the students. Each week, the designated student will lead a discussion of relevant topics/learning objectives needed to accomplish each week's task. This discussion will be followed by an instructor lead demonstration of a possible implementation of the weekly task. The second course period of each week will be a hands-on workshop where students will implement the weekly task themselves. Prereq: PQHS 431 and PQHS 432 .
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verified- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff