1 |
From FreEM to D'AlemBERT ; From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern French
|
|
|
|
In: Proceedings of the 13th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-03596653 ; Proceedings of the 13th Language Resources and Evaluation Conference, European Language Resources Association, Jun 2022, Marseille, France (2022)
|
|
BASE
|
|
Show details
|
|
2 |
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
|
|
|
|
In: https://hal.inria.fr/hal-03536361 ; 2022 (2022)
|
|
BASE
|
|
Show details
|
|
3 |
Ungoliant: An Optimized Pipeline for the Generation of a Very Large-Scale Multilingual Web Corpus
|
|
|
|
In: CMLC 2021 - 9th Workshop on Challenges in the Management of Large Corpora ; https://hal.inria.fr/hal-03301590 ; CMLC 2021 - 9th Workshop on Challenges in the Management of Large Corpora, Jul 2021, Limerick / Virtual, Ireland. ⟨10.14618/ids-pub-10468⟩ ; https://www.cl2021.org/ (2021)
|
|
BASE
|
|
Show details
|
|
4 |
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
|
|
|
|
In: https://hal.inria.fr/hal-03177623 ; 2021 (2021)
|
|
BASE
|
|
Show details
|
|
5 |
Establishing a New State-of-the-Art for French Named Entity Recognition
|
|
|
|
In: LREC 2020 - 12th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-02617950 ; LREC 2020 - 12th Language Resources and Evaluation Conference, May 2020, Marseille, France ; http://www.lrec-conf.org (2020)
|
|
BASE
|
|
Show details
|
|
6 |
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02889804 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, Canada. ⟨10.18653/v1/2020.acl-main.107⟩ (2020)
|
|
Abstract:
International audience ; We introduce the first treebank for a romanized user-generated content variety of Algerian, a North-African Arabic dialect known for its frequent usage of code-switching. Made of 1500 sentences, fully annotated in morpho-syntax and Universal Dependency syntax, with full translation at both the word and the sentence levels, this treebank is made freely available. It is supplemented with 50k unlabeled sentences collected from Common Crawl and web-crawled data using intensive data-mining techniques. Preliminary experiments demonstrate its usefulness for POS tagging and dependency parsing. We believe that what we present in this paper is useful beyond the low-resource language community. This is the first time that enough unlabeled and annotated data is provided for an emerging user-generated content dialectal language with rich morphology and code switching, making it an challenging test-bed for most recent NLP approaches.
|
|
Keyword:
[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]
|
|
URL: https://hal.inria.fr/hal-02889804/file/Building_an_Arabizi_Treebank__Tackling_Hell.pdf https://doi.org/10.18653/v1/2020.acl-main.107 https://hal.inria.fr/hal-02889804/document https://hal.inria.fr/hal-02889804
|
|
BASE
|
|
Hide details
|
|
7 |
CamemBERT: a Tasty French Language Model
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02889805 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States. ⟨10.18653/v1/2020.acl-main.645⟩ (2020)
|
|
BASE
|
|
Show details
|
|
8 |
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02863875 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States. ⟨10.18653/v1/2020.acl-main.156⟩ ; https://acl2020.org (2020)
|
|
BASE
|
|
Show details
|
|
9 |
French Contextualized Word-Embeddings with a sip of CaBeRnet: a New French Balanced Reference Corpus
|
|
|
|
In: CMLC-8 - 8th Workshop on the Challenges in the Management of Large Corpora ; https://hal.inria.fr/hal-02678358 ; CMLC-8 - 8th Workshop on the Challenges in the Management of Large Corpora, May 2020, Marseille, France ; https://lrec2020.lrec-conf.org/media/proceedings/Workshops/Books/CMLC-8book.pdf (2020)
|
|
BASE
|
|
Show details
|
|
10 |
Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
|
|
|
|
In: 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7) ; https://hal.inria.fr/hal-02148693 ; 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7), Jul 2019, Cardiff, United Kingdom. ⟨10.14618/IDS-PUB-9021⟩ (2019)
|
|
BASE
|
|
Show details
|
|
11 |
CamemBERT: a Tasty French Language Model
|
|
|
|
In: https://hal.inria.fr/hal-02445946 ; 2019 (2019)
|
|
BASE
|
|
Show details
|
|
12 |
Preparing the Dictionnaire Universel for Automatic Enrichment
|
|
|
|
In: 10th International Conference on Historical Lexicography and Lexicology (ICHLL) ; https://hal.inria.fr/hal-02131598 ; 10th International Conference on Historical Lexicography and Lexicology (ICHLL), Jun 2019, Leeuwarden, Netherlands ; https://easychair.org/smart-program/ICHLL-10/ (2019)
|
|
BASE
|
|
Show details
|
|
|
|