GEN 3898
Trustworthy Machine Learning: Building and evaluating agentic systems
Stanford University · UGRD · Fall 2026
Catalog description
This project-based course will introduce students to building and evaluating agentic AI applications powered by foundation models. The overriding theme of the course is that building an initial prototype AI system can often be completed easily but refining a prototype into something that is useful and reliable requires iterative improvement based on clear evaluation metrics. We will cover background in foundation models, prompting, and retrieval-augmented generation (RAG) before introducing full agentic AI architectures. For each architecture, the course will study methods for evaluation. Students will complete introductory homework assignments to become familiar with retrieval-augmented generation (RAG) and agentic AI. Students will then work in pairs or small teams to develop applications using agentic or other approaches and evaluate them by adapting evaluation methods presented in the class.Prerequisites: CS229 or similar introductory Python-based ML class; knowledge of deep learning such as CS230, CS231N; familiarity with ML frameworks in Python (scikit-learn, Keras) assumed.
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verified- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff