Catalogue search • Linguistik portal • Fachinformationsdienst (FID)

1	CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web ...
	The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021; Edunov, Sergey; Fan, angelafan@fb.com; Grave, Edouard; Joulin, Armand; Schwenk, Holger; Wenzek, Guillaume. - : Underline Science Inc., 2021
	Abstract: Read paper: https://www.aclanthology.org/2021.acl-long.507 Abstract: We show that margin-based bitext mining in a multilingual sentence space can be successfully scaled to operate on monolingual corpora of billions of sentences. We use 32 snapshots of a curated common crawl corpus (Wenzel et al, 2019) totaling 71 billion unique sentences. Using one unified approach for 90 languages, we were able to mine 10.8 billion parallel sentences, out of which only 2.9 billions are aligned with English. We illustrate the capability of our scalable mining system to create high quality training sets from one language to any other by training hundreds of different machine translation models and evaluating them on the many-to-many TED benchmark. Further, we evaluate on competitive translation benchmarks such as WMT and WAT. Using only mined bitext, we set a new state of the art for a single system on the WMT'19 test set for English-German/Russian/Chinese. In particular, our English/German and English/Russian systems ...
	URL: https://dx.doi.org/10.48448/z2vp-9188 https://underline.io/lecture/25726-ccmatrix-mining-billions-of-high-quality-parallel-sentences-on-the-web
	BASE
	Hide details

Search in the Catalogues and Directories