5 |
Controllable Sentence Simplification
|
|
|
|
In: LREC 2020 - 12th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-02678214 ; LREC 2020 - 12th Language Resources and Evaluation Conference, May 2020, Marseille, France ; http://www.lrec-conf.org/proceedings/lrec2020/index.html (2020)
|
|
BASE
|
|
Show details
|
|
6 |
CamemBERT: a Tasty French Language Model
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02889805 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States. ⟨10.18653/v1/2020.acl-main.645⟩ (2020)
|
|
BASE
|
|
Show details
|
|
7 |
French Contextualized Word-Embeddings with a sip of CaBeRnet: a New French Balanced Reference Corpus
|
|
|
|
In: CMLC-8 - 8th Workshop on the Challenges in the Management of Large Corpora ; https://hal.inria.fr/hal-02678358 ; CMLC-8 - 8th Workshop on the Challenges in the Management of Large Corpora, May 2020, Marseille, France ; https://lrec2020.lrec-conf.org/media/proceedings/Workshops/Books/CMLC-8book.pdf (2020)
|
|
Abstract:
International audience ; This paper describes and compares the impact of different types and size of training corpora on language models like ELMO. By asking the fundamental question of quality versus quantity we evaluate four French corpora for training on parsing scores, POS-tagging and named-entities recognition downstream tasks. The paper studies the relevance of a new corpus, CaBeRnet, featuring a representative range of language usage, including a balanced variety of genres (oral transcriptions, newspapers, popular magazines, technical reports, fiction, academic texts), in oral and written styles. We hypothesize that a linguistically representative and balanced corpora will allow the language model to be more efficient and representative of a given language and therefore yield better evaluation scores on different evaluation sets and tasks.
|
|
Keyword:
[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]; [SHS.LANGUE]Humanities and Social Sciences/Linguistics; Balanced French Corpus; BERT; ELMo; French; Language Models; NER; Parsing; Tagging
|
|
URL: https://hal.inria.fr/hal-02678358/file/LREC_Fabre_Ortiz.pdf https://hal.inria.fr/hal-02678358 https://hal.inria.fr/hal-02678358/document
|
|
BASE
|
|
Hide details
|
|
10 |
Controllable Sentence Simplification
|
|
|
|
In: https://hal.inria.fr/hal-02445874 ; 2019 (2019)
|
|
BASE
|
|
Show details
|
|
11 |
CamemBERT: a Tasty French Language Model
|
|
|
|
In: https://hal.inria.fr/hal-02445946 ; 2019 (2019)
|
|
BASE
|
|
Show details
|
|
12 |
Challenges of language change and variation: towards an extended treebank of Medieval French
|
|
|
|
In: TLT 2019 - 18th International Workshop on Treebanks and Linguistic Theories ; https://hal.inria.fr/hal-02272560 ; TLT 2019 - 18th International Workshop on Treebanks and Linguistic Theories, Aug 2019, Paris, France (2019)
|
|
BASE
|
|
Show details
|
|
13 |
Syntactic Parsing versus MWEs: What can fMRI signal tell us
|
|
|
|
In: PARSEME-FR 2019 consortium meeting ; https://hal.inria.fr/hal-02272288 ; PARSEME-FR 2019 consortium meeting, Jun 2019, Blois, France ; https://parsemefr.lis-lab.fr/doku.php?id=meeting-20190613 (2019)
|
|
BASE
|
|
Show details
|
|
14 |
Annotation tools for syntax
|
|
|
|
In: Rhapsodie: A Prosodic and Syntactic Treebank for Spoken French ; https://hal.inria.fr/hal-02450311 ; Rhapsodie: A Prosodic and Syntactic Treebank for Spoken French, John Benjamins, 2019, ⟨10.1075/scl.89.08ger⟩ (2019)
|
|
BASE
|
|
Show details
|
|
17 |
Reference-less Quality Estimation of Text Simplification Systems
|
|
|
|
In: 1st Workshop on Automatic Text Adaptation (ATA) ; https://hal.inria.fr/hal-01959054 ; 1st Workshop on Automatic Text Adaptation (ATA), Nov 2018, Tilburg, Netherlands ; https://www.ida.liu.se/~evere22/ATA-18/ (2018)
|
|
BASE
|
|
Show details
|
|
18 |
ELMoLex: Connecting ELMo and Lexicon features for Dependency Parsing
|
|
|
|
In: CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies ; https://hal.inria.fr/hal-01959045 ; CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, Oct 2018, Brussels, Belgium. ⟨10.18653/v1/K18-2023⟩ (2018)
|
|
BASE
|
|
Show details
|
|
19 |
Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer
|
|
|
|
In: Eleventh International Conference on Language Resources and Evaluation (LREC 2018) ; https://hal.inria.fr/hal-01798801 ; Eleventh International Conference on Language Resources and Evaluation (LREC 2018), May 2018, Miyazaki, Japan (2018)
|
|
BASE
|
|
Show details
|
|
20 |
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations
|
|
|
|
In: LREC 2018 - 11th edition of the Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-01744572 ; LREC 2018 - 11th edition of the Language Resources and Evaluation Conference, May 2018, Miyazaki, Japan ; http://lrec2018.lrec-conf.org/en/ (2018)
|
|
BASE
|
|
Show details
|
|
|
|