Logo image
Learning Accent Representation with Multi-Level VAE Towards Controllable Speech Synthesis
Conference proceeding

Learning Accent Representation with Multi-Level VAE Towards Controllable Speech Synthesis

Jan Melechovsky, Ambuj Mehrish, Dorien Herremans, Berrak Sisman and IEEE
2022 IEEE Spoken Language Technology Workshop (SLT), pp.928-935
09/01/2023

Abstract

Accent Conferences Controllable speech synthesis Disentanglement Fuses Measurement Multi-level Variational Autoencoder Proposals Speech synthesis Text-to-Speech
Accent is a crucial aspect of speech that helps define one's identity. We note that the state-of-the-art Text-to-Speech (TTS) systems can achieve high-quality generated voice, but still lack in terms of versatility and customizability. Moreover, they generally do not take into account accent, which is an important feature of speaking style. In this work, we utilize the concept of Multi-level VAE (ML-VAE) to build a control mechanism that aims to disentangle accent from a reference accented speaker; and to synthesize voices in different accents such as English, American, Irish, and Scottish. The proposed framework can also achieve high-quality accented voice generation for multi-speaker setup, which we believe is remarkable. We investigate the performance through objective metrics and conduct listening experiments for a subjective performance assessment. We showed that the proposed method achieves good performance for naturalness, speaker similarity, and accent similarity.

Metrics

1 Record Views

Details

Logo image