11 752

Speech Generation

Carnegie Mellon University · UGRD · Fall 2026

1 section
Add to a schedule

Catalog description

This course provides an in-depth study of speech synthesis and generative speech modeling, tracing the evolution from classical signal-processing and statistical approaches to modern neural and foundation-based methods. The course begins with the fundamentals of human speech production, which motivate early synthesis techniques such as concatenative and statistical parametric models, and then progresses to deep learning-based acoustic modeling, neural vocoding, and end-to-end speech generation architectures. Contemporary topics include the integration of large language models with speech synthesis, the use of discrete and continuous speech representations, and emerging paradigms for scalable and controllable speech generation. Beyond text-to-speech, the course covers a broad range of generative speech problems, including voice conversion, expressive speech synthesis, prosody generation and manipulation, speech editing, speech-to-speech translation, and spoken dialogue systems. Additional topics address disentangled representation learning, evaluation methodologies, speech databases, and security considerations such as spoofing attacks. Days of Week: Monday, Wednesday Time: 9:30 - 10:50 Enrollment is 30 Will you have a final exam? No, final project Letter grades or option to Pass/Fail? Letter grades Best regards, Carlos ________________________________ Professor, IEEE Fellow, ISCA Fellow Carnegie Mellon University School of Computer Science Language Technologies Institute

Sections

Current meeting, instructor, credit, and enrollment details

Updated 5 hours ago

001

Availability not recently verified
Class #carnegie_mellon-11752Fall 2026UGRD12 credits
Days & times
No scheduled meeting time
Meeting dates
Location
Instructor
Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?