ONE-SHOT VOICE CONVERSION BASED ON SPEAKER AWARE MODULE - Citegraph

Paper Info

Title
ONE-SHOT VOICE CONVERSION BASED ON SPEAKER AWARE MODULE

Abstract
Voice conversion (VC) is a task to convert the voice of speech while preserving its linguistic content. Although several methods have been proposed to enable VC with non-parallel data, it is still difficult to model the voice without a great number of data or an adaptive process. In this paper, we propose a speaker-aware voice conversion (SAVC) system realizing one-shot voice conversion without an adaptation stage. The SAVC utilizes a speaker aware module (SAM) to disentangle speaker embeddings. The SAM comprises a dynamic reference encoder, a static speaker knowledge block (SKB), and a multi-head attention layer. The reference encoder is used to compress a variable-length utterance to a fixed-length vector, the SKB is made up of pre-extraction x-vectors, and the multi-head attention layer is designed to generate weighted combined speaker embeddings. Subsequently, phonetic posteriorgrams (PPGs) as context encoding are concatenated with speaker embeddings and sent to the decoder module for generating acoustic features. Experimental results on the Aishell-1 corpus show that the proposed method can improve speaker similarity and converted utterances' speech quality.

Year	DOI	Venue
2021	10.1109/ICASSP39728.2021.9414081	2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021)
Keywords	DocType	Citations
speaker aware voice conversion, one-shot, phonetic posteriorgrams, x-vector	Conference	0
PageRank	References	Authors
0.34	0	6

Authors (6 rows)

Cited by (0 rows)

References (0 rows)

Name	Order	Citations	PageRank
Ying Zhang	1	0	0.68
Hao Che	2	0	0.68
Jie Li	3	44	2.43
Chenxing Li	4	14	6.76
Xiaorui Wang	5	19	6.13
Zhongyuan Wang	6	0	0.34

1