Options
2024
Conference Paper
Title
PAD-VC: A Prosody-Aware Decoder for Any-to-Few Voice Conversion
Abstract
Voice Conversion (VC) generates synthetic speech from a source speaker recording, preserving linguistic information and applying the voice characteristics of a target speaker. In this paper, we propose PAD-VC, a prosody-aware VC model based on the decoder part of the ForwardTacotron architecture. We train PAD-VC with prosody-related features such as pitch, energy, and voicing confidence and augment those with linguistic features derived from a phoneme posteriorgram representation of the source utterance. This way, we can handle both phonemic information and framewise supra-segmental features. During inference time, the source speaker's prosody features are modified to match the prosody statistics of the target speaker. We show that PAD-VC outperforms ForwardTacotron in prosody-cloning for unseen source utterances, achieving higher similarity and naturalness.
Author(s)
Mainwork
2024 18th International Workshop on Acoustic Signal Enhancement Iwaenc 2024 Proceedings
Funder
Indo-German Science and Technology Centre
Conference
18th International Workshop on Acoustic Signal Enhancement, IWAENC 2024