LING 696G: Advanced Speech Technology

Neural techniques for speech synthesis (TTS) and automatic speech recognition (ASR/STT). Hands-on implementation of neural speech systems on HPC.

Instructor: Daniel Brenner

Term: Spring

Location: Online

Course Overview

This course introduces neural techniques for speech synthesis and automatic speech recognition. Starting with overviews of TTS, ASR, and neural networks, the course moves into specific neural models. Grading is based on neural systems that students implement.

Topics

  • Neural networks — overview for speech applications
  • Text-to-speech (TTS) — neural synthesis systems
  • Speech-to-text (STT) — neural recognition systems
  • Transfer learning — adapting pretrained models

Requirements

Students implement TTS or STT systems on HPC, culminating in a final project that develops and iteratively improves a neural speech system.

Prerequisites

Familiarity with content of LING 578 recommended. Python programming required.

Textbook

Hammond, Michael (2025) Speech Technology (draft textbook, available on D2L).