Department of Information and Communications Engineering

Speech Synthesis

Speech synthesis research group studies deep generative models and differentiable signal processing methods applied to speech and audio. Further research interests include voice cloning from limited data, deepfake detection, countermeasures, and watermarking.
graph

Speaking machines have long been a central research interest in speech processing and machine learning. Speech synthesis is a key component in, for example, conversational agents, personal digital assistants, audiobook readers, and assistive devices including screen readers and voice prostheses. Modern speech synthesis methods achieve close to human-level naturalness by using deep generative models, such as WaveNet, GANs, Diffusion models and Transformer language models.

Current technical challenges in speech synthesis include efficiency, control, and interpretability. State-of-the-art relies on large neural network models, which are computationally expensive black-boxes. Aalto Speech Synthesis Group research combines classic digital signal processing methods with differentiable computing for efficient and interpretable neural synthesis.

Instant voice cloning is another trend in speech synthesis technology. A voice cloning system can be adapted to a new voice from just a few seconds of audio, which opens many exciting applications, but also presents a pressing set of challenges for deepfake detection. Building responsible synthesis with watermarking is a current research topic in the Aalto Speech Synthesis Group.

The Aalto Speech Synthesis Group has ongoing collaboration related to the above topics with KTH Royal Institute of Technology, Sweden; and National Institute of Informatics (NII), Japan.

Current research topics

  • Generative models for speech synthesis: GANs, WaveNets, diffusion models, discrete representation learning from audio, language models for sound
  • Differentiable DSP: digital signal processing as building blocks for efficient neural synthesis systems
  • Watermarking generative models; Speech deepfake detection, countermeasures, and awareness

The Speech Synthesis Research Group is led by professor Lauri Juvela

Group members

Latest publications

Aliasing-Free Neural Audio Synthesis

Yicheng Gu, Junan Zhang, Chaoren Wang, Jerry Li, Zhizheng Wu, Lauri Juvela 2026 IEEE Transactions on Audio, Speech and Language Processing

Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN

Yicheng Gu, Chaoren Wang, Zhizheng Wu, Lauri Juvela 2025 Proceedings of the Interspeech

SOLID STATE BUS-COMP: A LARGE-SCALE AND DIVERSE DATASET FOR DYNAMIC RANGE COMPRESSOR VIRTUAL ANALOG MODELING

Yicheng Gu, Runsong Zhang, Lauri Juvela, Zhizheng Wu 2025 Proceedings of the 28th International Conference on Digital Audio Effects

Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis

Lauri Juvela, Xin Wang 2025 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Proceedings

Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams

Zirui Li, Lauri Juvela, Mikko Kurimo 2025 13th edition of the Speech Synthesis Workshop

Unsupervised Estimation of Nonlinear Audio Effects: Comparing Diffusion-based and Adversarial Approaches

Eloi Moliner Juanpere, Michal Švento, Alec Wright, Lauri Juvela, Pavel Rajmic, Vesa Välimäki 2025 Proceedings of the 28th International Conference on Digital Audio Effects

Estimation and Restoration of Unknown Nonlinear Distortion Using Diffusion

Michal Švento, Eloi Moliner Juanpere, Lauri Juvela, Alec Wright, Vesa Välimäki 2025 Journal of the Audio Engineering Society

Open-Amp: Synthetic Data Framework for Audio Effect Foundation Models

Alec Wright, Alistair Carson, Lauri Juvela 2025 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Proceedings

KLANN: Linearising Long-Term Dynamics in Nonlinear Audio Effects Using Koopman Networks

Ville Huhtala, Lauri Juvela, Sebastian J. Schlecht 2024 IEEE Signal Processing Letters

DDSP-based Neural Waveform Synthesis of Polyphinic Guitar Performance From String-Wise Midi Input

Nicolas Jonason, Xin Wang, Erica Cooper, Lauri Juvela, Bob L.T. Sturm, Junichi Yamagishi 2024 Proceedings of the 27th International Conference on Digital Audio Effects (DAFx24)
More information on our research in the Aalto research portal.
Research portal
  • Updated:
  • Published:
Share
URL copied!