DE eng

Search in the Catalogues and Directories

Hits 1 – 14 of 14

1
FlauBERT: Unsupervised Language Model Pre-training for French
In: Proceedings of the 12th Language Resources and Evaluation Conference ; LREC ; https://hal.archives-ouvertes.fr/hal-02890258 ; LREC, 2020, Marseille, France (2020)
Abstract: International audience ; Language models have become a key step to achieve state-of-the art results in many different Natural Language Processing (NLP) tasks. Leveraging the huge amount of unlabeled texts nowadays available, they provide an efficient way to pre-train continuous word representations that can be fine-tuned for a downstream task, along with their contextualization at the sentence level. This has been widely demonstrated for English using contextualized representations (Dai and Le, 2015; Peters et al., 2018; Howard and Ruder, 2018; Radford et al., 2018; Devlin et al., 2019; Yang et al., 2019b). In this paper, we introduce and share FlauBERT, a model learned on a very large and heterogeneous French corpus. Models of different sizes are trained using the new CNRS (French National Centre for Scientific Research) Jean Zay supercomputer. We apply our French language models to diverse NLP tasks (text classification, paraphrasing, natural language inference, parsing, word sense disambiguation) and show that most of the time they outperform other pre-training approaches. Different versions of FlauBERT as well as a unified evaluation protocol for the downstream tasks, called FLUE (French Language Understanding Evaluation), are shared to the research community for further reproducible experiments in French NLP.
Keyword: [INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]; BERT; FlauBERT; FLUE; French; language model; natural language inference; NLP benchmark; paraphrase; parsing; pre-training; text classification; Transformer; word sense disambiguation
URL: https://hal.archives-ouvertes.fr/hal-02890258/file/Flaubert.pdf
https://hal.archives-ouvertes.fr/hal-02890258
https://hal.archives-ouvertes.fr/hal-02890258/document
BASE
Hide details
2
FlauBERT : Unsupervised Language Model Pre-training for French ; FlauBERT : des modèles de langue contextualisés pré-entraînés pour le français
In: Actes de la 6e conférence conjointe Journées d'Études sur la Parole (JEP, 33e édition), Traitement Automatique des Langues Naturelles (TALN, 27e édition), Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL, 22e édition). Volume 2 : Traitement Automatique des Langues Naturelles ; 6e conférence conjointe Journées d'Études sur la Parole (JEP, 33e édition), Traitement Automatique des Langues Naturelles (TALN, 27e édition), Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL, 22e édition). Volume 2 : Traitement Automatique des Langues Naturelles ; https://hal.archives-ouvertes.fr/hal-02784776 ; 6e conférence conjointe Journées d'Études sur la Parole (JEP, 33e édition), Traitement Automatique des Langues Naturelles (TALN, 27e édition), Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL, 22e édition). Volume 2 : Traitement Automatique des Langues Naturelles, Jun 2020, Nancy, France. pp.268-278 (2020)
BASE
Show details
3
STEM-ECR-v1.0 ...
BASE
Show details
4
The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources ...
D'Souza, Jennifer; Hoppe, Anett; Brack, Arthur. - : Paris : European Language Resources Association, 2020
BASE
Show details
5
Общий взгляд на криминалистическое следоведение ... : GENERAL VIEW ON FORENSIC INVESTIGATION ...
Ищенко Евгений Петрович. - : Вестник Восточно-Сибирского института МВД России, 2019
BASE
Show details
6
A novel framework for biomedical entity sense induction
In: ISSN: 1532-0464 ; EISSN: 1532-0480 ; Journal of Biomedical Informatics ; https://hal-lirmm.ccsd.cnrs.fr/lirmm-01851988 ; Journal of Biomedical Informatics, Elsevier, 2018, 84, pp.31-41. ⟨10.1016/j.jbi.2018.06.007⟩ (2018)
BASE
Show details
7
Cross-domain polarity classification using a knowledge-enhanced meta-classifier
BASE
Show details
8
Senso Comune as a Knowledge Base of Italian language: The Resource and its Development
In: Proceedings of the First Italian Conference on Computational Linguistics CLiC-it 2014 ; First Italian Conference on Computational Linguistics - CLiC-it 2014 ; https://hal.archives-ouvertes.fr/hal-01134621 ; First Italian Conference on Computational Linguistics - CLiC-it 2014, Dec 2014, Pisa, Italy. pp. 93-97 (2014)
BASE
Show details
9
Applying Lexical Semantics to Improve Text Classification
In: http://web2py.iiit.ac.in/publications/default/download/inproceedings.pdf.9ecb6867-0fb0-48a5-8020-0310468d3275.pdf (2008)
BASE
Show details
10
Hybrid Methods for Acquisition of Lexical Information: the Case for Verbs
In: http://rave.ohiolink.edu/etdc/view?acc_num=osu1228259857 (2008)
BASE
Show details
11
Desambiguación léxica mediante marcas de especificidad
Montoyo, Andres. - : Sociedad Española para el Procesamiento del Lenguaje Natural, 2002
BASE
Show details
12
Proceedings of the Twenty-Fourth Conference on Computational Linguistics and Speech Processing (ROCLING 2012) Associating Collocations with WordNet Senses Using Hybrid Models �
In: http://wing.comp.nus.edu.sg/~antho/O/O12/O12-1006.pdf
BASE
Show details
13
Ontology-Supported Text Classification Based on Cross-Lingual Word Sense Disambiguation
In: http://www.racai.ro/~tufis/papers/Tufis-Koeva-wilf2007.pdf
BASE
Show details
14
Many tasks in Natural Language Processing – Word
In: http://www.plwordnet.pwr.wroc.pl/main/content/files/publications/ltc07-125-piasecki.pdf
BASE
Show details

Catalogues
0
0
0
0
0
0
0
Bibliographies
0
0
0
0
0
0
0
0
0
Linked Open Data catalogues
0
Online resources
0
0
0
0
Open access documents
14
0
0
0
0
© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern