1 |
Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization ...
|
|
|
|
BASE
|
|
Show details
|
|
2 |
Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation ...
|
|
|
|
Abstract:
A conventional approach to improving the performance of end-to-end speech translation (E2E-ST) models is to leverage the source transcription via pre-training and joint training with automatic speech recognition (ASR) and neural machine translation (NMT) tasks. However, since the input modalities are different, it is difficult to leverage source language text successfully. In this work, we focus on sequence-level knowledge distillation (SeqKD) from external text-based NMT models. To leverage the full potential of the source language information, we propose backward SeqKD, SeqKD from a target-to-source backward NMT model. To this end, we train a bilingual E2E-ST model to predict paraphrased transcriptions as an auxiliary task with a single decoder. The paraphrases are generated from the translations in bitext via back-translation. We further propose bidirectional SeqKD in which SeqKD from both forward and backward NMT models is combined. Experimental evaluations on both autoregressive and non-autoregressive ... : Accepted at NAACL-HLT 2021 (short paper) ...
|
|
Keyword:
Audio and Speech Processing eess.AS; Computation and Language cs.CL; FOS Computer and information sciences; FOS Electrical engineering, electronic engineering, information engineering; Sound cs.SD
|
|
URL: https://arxiv.org/abs/2104.06457 https://dx.doi.org/10.48550/arxiv.2104.06457
|
|
BASE
|
|
Hide details
|
|
3 |
Self-Guided Curriculum Learning for Neural Machine Translation ...
|
|
|
|
BASE
|
|
Show details
|
|
4 |
Arabic Speech Recognition by End-to-End, Modular Systems and Human ...
|
|
|
|
BASE
|
|
Show details
|
|
5 |
Leveraging End-to-End ASR for Endangered Language Documentation: An Empirical Study on Yoloxóchitl Mixtec ...
|
|
|
|
BASE
|
|
Show details
|
|
7 |
Leveraging Pre-trained Language Model for Speech Sentiment Analysis ...
|
|
|
|
BASE
|
|
Show details
|
|
8 |
End-to-end ASR to jointly predict transcriptions and linguistic annotations ...
|
|
|
|
BASE
|
|
Show details
|
|
9 |
Differentiable Allophone Graphs for Language-Universal Speech Recognition ...
|
|
|
|
BASE
|
|
Show details
|
|
10 |
Speech Representation Learning Combining Conformer CPC with Deep Cluster for the ZeroSpeech Challenge 2021 ...
|
|
|
|
BASE
|
|
Show details
|
|
11 |
CHiME-6 Challenge: Tackling multispeaker speech recognition for unsegmented recordings
|
|
|
|
In: CHiME 2020 - 6th International Workshop on Speech Processing in Everyday Environments ; https://hal.inria.fr/hal-02546993 ; CHiME 2020 - 6th International Workshop on Speech Processing in Everyday Environments, May 2020, Barcelona / Virtual, Spain (2020)
|
|
BASE
|
|
Show details
|
|
14 |
A Comparative Study on Transformer vs RNN in Speech Applications ...
|
|
|
|
BASE
|
|
Show details
|
|
16 |
Towards Online End-to-end Transformer Automatic Speech Recognition ...
|
|
|
|
BASE
|
|
Show details
|
|
18 |
The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines
|
|
|
|
In: Interspeech 2018 - 19th Annual Conference of the International Speech Communication Association ; https://hal.inria.fr/hal-01744021 ; Interspeech 2018 - 19th Annual Conference of the International Speech Communication Association, Sep 2018, Hyderabad, India (2018)
|
|
BASE
|
|
Show details
|
|
19 |
Analysis of Multilingual Sequence-to-Sequence speech recognition systems ...
|
|
|
|
BASE
|
|
Show details
|
|
20 |
Language model integration based on memory control for sequence to sequence speech recognition ...
|
|
|
|
BASE
|
|
Show details
|
|
|
|