LING 696G: Advanced Speech Technology
Neural techniques for speech synthesis (TTS) and automatic speech recognition (ASR/STT). Hands-on implementation of neural speech systems on HPC.
Instructor: Daniel Brenner
Term: Spring
Location: Online
Course Overview
This course introduces neural techniques for speech synthesis and automatic speech recognition. Starting with overviews of TTS, ASR, and neural networks, the course moves into specific neural models. Grading is based on neural systems that students implement.
Topics
- Neural networks — overview for speech applications
- Text-to-speech (TTS) — neural synthesis systems
- Speech-to-text (STT) — neural recognition systems
- Transfer learning — adapting pretrained models
Requirements
Students implement TTS or STT systems on HPC, culminating in a final project that develops and iteratively improves a neural speech system.
Prerequisites
Familiarity with content of LING 578 recommended. Python programming required.
Textbook
Hammond, Michael (2025) Speech Technology (draft textbook, available on D2L).