GEN 3780
Mechanistic Interpretability
Stanford University · UGRD · Fall 2026
1 section
Catalog description
What is the internal structure of modern neural networks and how can we study it? This course provides a broad and deep introduction to interpretability, the subfield of machine learning concerned with understanding precisely how models process information and why they produce the outputs they do. We will cover topics such as probing, steering, causal abstraction, and sparse autoencoders, with a particular emphasis on causal methods and large language models. The course will include guest lectures from leading interpretability labs across academia and industry.
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verifiedClass #stanford-3780Fall 2026UGRD3 credits
- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?