Abstract
This paper proposes a method of disentangling accent features from speaker representations during voice conversion through auto-encoder and adversarial training techniques. Non-parallel data can be used during training, addressing cases of limited data. To generate samples, the three components of linguistic content, accent representations and speaker representations are extracted from acoustic features and disentangled. To ensure that relevant features are captured by the encoders, an auto-encoder training framework was employed where decoders make use of latent variables extracted by encoders in an attempt to reconstruct the input sample. Similar to the work by J.-X. Zhang, Ling, and Dai (2019), linguistic content in acoustic features is identified through the use of ground truth transcripts. To eliminate speaker-related features from the linguistic content representations and accent-related features from the speaker representations, adversarial training was employed. A two-stage training methodology was employed to train the model - the model is first pre-trained on a large dataset with multiple native speakers of different accents to ensure that linguistic content and speaker representations can be disentangled, after which, the model is then fine-tuned on a relatively smaller dataset of multiple non-native speakers to allow the model to learn differences between various non-native accents and disentangle accent characteristics from speaker representations. Based on listening tests, the proposed model scored higher than the baseline for naturalness and speaker similarity, and above average for accent similarity.