GEN 3898

Trustworthy Machine Learning: Building and evaluating agentic systems

Stanford University · UGRD · Fall 2026

1 section
Add to a schedule

Catalog description

This project-based course will introduce students to building and evaluating agentic AI applications powered by foundation models. The overriding theme of the course is that building an initial prototype AI system can often be completed easily but refining a prototype into something that is useful and reliable requires iterative improvement based on clear evaluation metrics. We will cover background in foundation models, prompting, and retrieval-augmented generation (RAG) before introducing full agentic AI architectures. For each architecture, the course will study methods for evaluation. Students will complete introductory homework assignments to become familiar with retrieval-augmented generation (RAG) and agentic AI. Students will then work in pairs or small teams to develop applications using agentic or other approaches and evaluate them by adapting evaluation methods presented in the class.Prerequisites: CS229 or similar introductory Python-based ML class; knowledge of deep learning such as CS230, CS231N; familiarity with ML frameworks in Python (scikit-learn, Keras) assumed.

Sections

Current meeting, instructor, credit, and enrollment details

Updated 4 hours ago

001

Availability not recently verified
Class #stanford-3898Fall 2026UGRD3 credits
Days & times
No scheduled meeting time
Meeting dates
Location
Instructor
Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?