Home Catalogue search

eng

Refine your search:

Search in the Catalogues and Directories






	Sort by
Simple Search

Hits 1 – 4 of 4

1	Learning Speaker Embedding from Text-to-Speech ...
	Cho, Jaejin; Zelasko, Piotr; Villalba, Jesus; Watanabe, Shinji; Dehak, Najim. - : arXiv, 2020
	Abstract: Zero-shot multi-speaker Text-to-Speech (TTS) generates target speaker voices given an input text and the corresponding speaker embedding. In this work, we investigate the effectiveness of the TTS reconstruction objective to improve representation learning for speaker verification. We jointly trained end-to-end Tacotron 2 TTS and speaker embedding networks in a self-supervised fashion. We hypothesize that the embeddings will contain minimal phonetic information since the TTS decoder will obtain that information from the textual input. TTS reconstruction can also be combined with speaker classification to enhance these embeddings further. Once trained, the speaker encoder computes representations for the speaker verification task, while the rest of the TTS blocks are discarded. We investigated training TTS from either manual or ASR-generated transcripts. The latter allows us to train embeddings on datasets without manual transcripts. We compared ASR transcripts and Kaldi phone alignments as TTS inputs, showing ...
	Keyword: Audio and Speech Processing eess.AS; FOS Computer and information sciences; FOS Electrical engineering, electronic engineering, information engineering; Machine Learning cs.LG; Sound cs.SD
	URL: https://arxiv.org/abs/2010.11221 https://dx.doi.org/10.48550/arxiv.2010.11221
	BASE
	Hide details

2	Language model integration based on memory control for sequence to sequence speech recognition ...
	Cho, Jaejin; Watanabe, Shinji; Hori, Takaaki. - : arXiv, 2018
	BASE
	Show details

3	Transfer learning of language-independent end-to-end ASR with language model fusion ...
	Inaguma, Hirofumi; Cho, Jaejin; Baskar, Murali Karthick. - : arXiv, 2018
	BASE
	Show details

4	Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling ...
	Cho, Jaejin; Baskar, Murali Karthick; Li, Ruizhi. - : arXiv, 2018
	BASE
	Show details

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern