11 801

Quantitative Evaluation of Language Technologies

Carnegie Mellon University · UGRD · Fall 2026

1 section
Add to a schedule

Catalog description

Evaluating NLP models for properties like quality, fluency, safety, or revenue generation is a fundamental part of model development in research and engineering contexts. This course will present fundamental principles of evaluation, focusing on both offline contexts such as benchmark datasets and online environments such as production systems. Material will be organized into two parts. The first part will focus on measurement of system decisions, covering the principles of metric design (i.e., what makes a metric appropriate for a given NLP task), data elicitation (i.e., how can we gather data associated with what we are trying to measure), and modeling system properties (i.e., how can we formally model the properties like quality). The second part will focus on comparing systems given a set of measured values. We will cover dataset construction and hypothesis testing in both offline and online contexts. The course will include lectures from external researchers and practitioners using evaluation techniques to audit systems, design new benchmarks, and deploy production models.

Sections

Current meeting, instructor, credit, and enrollment details

Updated 4 hours ago

001

Availability not recently verified
Class #carnegie_mellon-11801Fall 2026UGRD12 credits
Days & times
No scheduled meeting time
Meeting dates
Location
Instructor
Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?