LING-GA 1012
Large Language Models: Evaluation and Applications
New York University · UGRD · Fall 2026
Catalog description
Contemporary language models can conduct conversations with users and perform tasks that go far beyond their historical role of predicting the next word. As these abilities are not explicitly engineered into the architecture, but rather emerge from large-scale training, it is often unclear how the models accomplish those tasks. In this course, we will survey the tasks, methods and metrics that are used to evaluate and understand language models, and to compare these models to humans. Some of the areas of evaluation we will discuss include factuality, pragmatics, reasoning, fairness and safety. We will also discuss interpretability methods that aim to explain the internal mechanisms that underlie the models' behavior. A major aim of the course is to prepare students to do original research in this area, culminating with a substantial final project.
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verified- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff