GEN 3780

Mechanistic Interpretability

Stanford University · UGRD · Fall 2026

1 section
Add to a schedule

Catalog description

What is the internal structure of modern neural networks and how can we study it? This course provides a broad and deep introduction to interpretability, the subfield of machine learning concerned with understanding precisely how models process information and why they produce the outputs they do. We will cover topics such as probing, steering, causal abstraction, and sparse autoencoders, with a particular emphasis on causal methods and large language models. The course will include guest lectures from leading interpretability labs across academia and industry.

Sections

Current meeting, instructor, credit, and enrollment details

Updated 3 hours ago

001

Availability not recently verified
Class #stanford-3780Fall 2026UGRD3 credits
Days & times
No scheduled meeting time
Meeting dates
Location
Instructor
Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?