11 752
Speech Generation
Carnegie Mellon University · UGRD · Fall 2026
Catalog description
This course provides an in-depth study of speech synthesis and generative speech modeling, tracing the evolution from classical signal-processing and statistical approaches to modern neural and foundation-based methods. The course begins with the fundamentals of human speech production, which motivate early synthesis techniques such as concatenative and statistical parametric models, and then progresses to deep learning-based acoustic modeling, neural vocoding, and end-to-end speech generation architectures. Contemporary topics include the integration of large language models with speech synthesis, the use of discrete and continuous speech representations, and emerging paradigms for scalable and controllable speech generation. Beyond text-to-speech, the course covers a broad range of generative speech problems, including voice conversion, expressive speech synthesis, prosody generation and manipulation, speech editing, speech-to-speech translation, and spoken dialogue systems. Additional topics address disentangled representation learning, evaluation methodologies, speech databases, and security considerations such as spoofing attacks. Days of Week: Monday, Wednesday Time: 9:30 - 10:50 Enrollment is 30 Will you have a final exam? No, final project Letter grades or option to Pass/Fail? Letter grades Best regards, Carlos ________________________________ Professor, IEEE Fellow, ISCA Fellow Carnegie Mellon University School of Computer Science Language Technologies Institute
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verified- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff