BMIN 5220
Natural Language Processing for Health
University of Pennsylvania · UGRD · Fall 2026
Catalog description
The growing volume of unstructured health-related data presents unparalleled challenges and opportunities for informaticians, clinicians, epidemiologists and other public health researchers that seek to mine the rich information "locked" within free-texts. Clinical records, social media, published literature, transcribed text, among other textual sources are designed for human eyes, but not necessarily for automatic processing. In this class, we will survey the most recent natural language processing methods used for identifying and classifying information present in these sources. The class provides learning of health language processing – that is, the fundamental principles and methods of both natural language processing and machine learning and how they are currently applied in the biomedical domain. The class will focus on real problems in the context of health research where data are inherently biased, e.g., noisy, missing, or extremely imbalanced. Methods for addressing these biases, such as text normalization, rules-based systems, machine learning (supervised, unsupervised, active learning), deep learning, and large language models will be discussed. In-class lectures will be most often taught using Jupyter notebooks and guest speakers presenting how an NLP/ML method was used to solve a driving biomedical use case. This course requires proficiency in python programming and machine learning. NOTE: Non-majors need permission from the instructor.
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verified- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff