Home Catalogue search

eng

Refine your search:

Search in the Catalogues and Directories






	Sort by
Simple Search

Page: 1 2 3 4 5...36

Hits 1 – 20 of 703

1	Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech Translation ...
	Fukuda, Ryo; Sudoh, Katsuhito; Nakamura, Satoshi. - : arXiv, 2022
	BASE
	Show details

2	A New Amharic Speech Emotion Dataset and Classification Benchmark ...
	Retta, Ephrem A.; Almekhlafi, Eiad; Sutcliffe, Richard; Mhamed, Mustafa; Ali, Haider; Feng, Jun. - : arXiv, 2022
	Abstract: In this paper we present the Amharic Speech Emotion Dataset (ASED), which covers four dialects (Gojjam, Wollo, Shewa and Gonder) and five different emotions (neutral, fearful, happy, sad and angry). We believe it is the first Speech Emotion Recognition (SER) dataset for the Amharic language. 65 volunteer participants, all native speakers, recorded 2,474 sound samples, two to four seconds in length. Eight judges assigned emotions to the samples with high agreement level (Fleiss kappa = 0.8). The resulting dataset is freely available for download. Next, we developed a four-layer variant of the well-known VGG model which we call VGGb. Three experiments were then carried out using VGGb for SER, using ASED. First, we investigated whether Mel-spectrogram features or Mel-frequency Cepstral coefficient (MFCC) features work best for Amharic. This was done by training two VGGb SER models on ASED, one using Mel-spectrograms and the other using MFCC. Four forms of training were tried, standard cross-validation, and ... : 16 pages, 12 tables, 6 figures ...
	Keyword: Audio and Speech Processing eess.AS; Computation and Language cs.CL; FOS Computer and information sciences; FOS Electrical engineering, electronic engineering, information engineering; I.2.7; Sound cs.SD
	URL: https://dx.doi.org/10.48550/arxiv.2201.02710 https://arxiv.org/abs/2201.02710
	BASE
	Hide details

3	Lahjoita puhetta -- a large-scale corpus of spoken Finnish with some benchmarks ...
	Moisio, Anssi; Porjazovski, Dejan; Rouhe, Aku. - : arXiv, 2022
	BASE
	Show details

4	The Norwegian Parliamentary Speech Corpus ...
	Solberg, Per Erik; Ortiz, Pablo. - : arXiv, 2022
	BASE
	Show details

5	Subspace-based Representation and Learning for Phonotactic Spoken Language Recognition ...
	Lee, Hung-Shin; Tsao, Yu; Jeng, Shyh-Kang. - : arXiv, 2022
	BASE
	Show details

6	Automatic Dialect Density Estimation for African American English ...
	Johnson, Alexander; Everson, Kevin; Ravi, Vijay. - : arXiv, 2022
	BASE
	Show details

7	End-to-end contextual asr based on posterior distribution adaptation for hybrid ctc/attention system ...
	Zhang, Zhengyi; Zhou, Pan. - : arXiv, 2022
	BASE
	Show details

8	Towards Contextual Spelling Correction for Customization of End-to-end Speech Recognition Systems ...
	Wang, Xiaoqiang; Liu, Yanqing; Li, Jinyu. - : arXiv, 2022
	BASE
	Show details

9	SHAS: Approaching optimal Segmentation for End-to-End Speech Translation ...
	Tsiamas, Ioannis; Gállego, Gerard I.; Fonollosa, José A. R.. - : arXiv, 2022
	BASE
	Show details

10	DanFEVER: claim verification dataset for Danish ...
	Nørregaard, Jeppe; Derczynski, Leon. - : figshare, 2022
	BASE
	Show details

11	DanFEVER: claim verification dataset for Danish ...
	Nørregaard, Jeppe; Derczynski, Leon. - : figshare, 2022
	BASE
	Show details

12	Repeat after me: Self-supervised learning of acoustic-to-articulatory mapping by vocal imitation ...
	Georges, Marc-Antoine; Diard, Julien; Girin, Laurent. - : arXiv, 2022
	BASE
	Show details

13	Synthesizing Dysarthric Speech Using Multi-talker TTS for Dysarthric Speech Recognition ...
	Soleymanpour, Mohammad; Johnson, Michael T.; Soleymanpour, Rahim. - : arXiv, 2022
	BASE
	Show details

14	Giant Pigeon and Small Person: Prompting Visually Grounded Models about the Size of Objects ...
	Zhang, Yi. - : Purdue University Graduate School, 2022
	BASE
	Show details

15	Giant Pigeon and Small Person: Prompting Visually Grounded Models about the Size of Objects ...
	Zhang, Yi. - : Purdue University Graduate School, 2022
	BASE
	Show details

16	A Hierarchical Model for Spoken Language Recognition ...
	Ferrer, Luciana; Castan, Diego; McLaren, Mitchell. - : arXiv, 2022
	BASE
	Show details

17	Hierarchical Softmax for End-to-End Low-resource Multilingual Speech Recognition ...
	Liu, Qianying; Yang, Yuhang; Gong, Zhuo. - : arXiv, 2022
	BASE
	Show details

18	Cross-lingual Self-Supervised Speech Representations for Improved Dysarthric Speech Recognition ...
	Hernandez, Abner; Pérez-Toro, Paula Andrea; Nöth, Elmar. - : arXiv, 2022
	BASE
	Show details

19	Multilingual Simultaneous Speech Translation ...
	Subramanya, Shashank; Niehues, Jan. - : arXiv, 2022
	BASE
	Show details

20	Code-Switching Text Augmentation for Multilingual Speech Processing ...
	Hussein, Amir; Chowdhury, Shammur Absar; Abdelali, Ahmed. - : arXiv, 2022
	BASE
	Show details

Page: 1 2 3 4 5...36

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern