1 |
Establishing a New State-of-the-Art for French Named Entity Recognition
|
|
|
|
In: LREC 2020 - 12th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-02617950 ; LREC 2020 - 12th Language Resources and Evaluation Conference, May 2020, Marseille, France ; http://www.lrec-conf.org (2020)
|
|
BASE
|
|
Show details
|
|
2 |
OFrLex: A Computational Morphological and Syntactic Lexicon for Old French
|
|
|
|
In: LREC 2020 - 12th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-02677957 ; LREC 2020 - 12th Language Resources and Evaluation Conference, May 2020, Marseille, France. 3217-3225 (updated version) (2020)
|
|
BASE
|
|
Show details
|
|
3 |
Controllable Sentence Simplification
|
|
|
|
In: LREC 2020 - 12th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-02678214 ; LREC 2020 - 12th Language Resources and Evaluation Conference, May 2020, Marseille, France ; http://www.lrec-conf.org/proceedings/lrec2020/index.html (2020)
|
|
BASE
|
|
Show details
|
|
4 |
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02889804 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, Canada. ⟨10.18653/v1/2020.acl-main.107⟩ (2020)
|
|
BASE
|
|
Show details
|
|
5 |
CamemBERT: a Tasty French Language Model
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02889805 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States. ⟨10.18653/v1/2020.acl-main.645⟩ (2020)
|
|
BASE
|
|
Show details
|
|
6 |
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02863875 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States. ⟨10.18653/v1/2020.acl-main.156⟩ ; https://acl2020.org (2020)
|
|
BASE
|
|
Show details
|
|
7 |
Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB 2.0
|
|
|
|
In: LREC 2020 - 12th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-02678100 ; LREC 2020 - 12th Language Resources and Evaluation Conference, May 2020, Marseille, France (2020)
|
|
Abstract:
Due to COVID19 pandemic, the 12th edition is cancelled. The LREC 2020 Proceedings are available at http://www.lrec-conf.org/proceedings/lrec2020/index.html ; International audience ; Diachronic lexical information was mostly used in its natural field, historical linguistics, until recently, when promising but not yet conclusive applications to low resource languages machine translation started extending its usage to NLP. There is therefore a new need for fine-grained, large-coverage and accurate etymological lexical resources. In this paper, we propose a set of guidelines to generate such resources, for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation. To illustrate the guidelines, we introduce EtymDB 2.0, an etymological database automatically generated from the Wiktionary, which contains 1.8 million lexemes, linked by more than 700,000 fine-grained etymological relations, across 2,536 living and dead languages. We also introduce use cases for which EtymDB 2.0 could represent a key resource, such as phylogenetic tree generation, low resource machine translation and medieval languages study.
|
|
Keyword:
[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]; [SHS.LANGUE]Humanities and Social Sciences/Linguistics; Etymological lexicon; Language Resource Life-cycle; Lexical Resource Development; Methodology
|
|
URL: https://hal.inria.fr/hal-02678100 https://hal.inria.fr/hal-02678100/document https://hal.inria.fr/hal-02678100/file/Updating_and_correcting_an_etymological_database___LREC-4.pdf
|
|
BASE
|
|
Hide details
|
|
8 |
French Contextualized Word-Embeddings with a sip of CaBeRnet: a New French Balanced Reference Corpus
|
|
|
|
In: CMLC-8 - 8th Workshop on the Challenges in the Management of Large Corpora ; https://hal.inria.fr/hal-02678358 ; CMLC-8 - 8th Workshop on the Challenges in the Management of Large Corpora, May 2020, Marseille, France ; https://lrec2020.lrec-conf.org/media/proceedings/Workshops/Books/CMLC-8book.pdf (2020)
|
|
BASE
|
|
Show details
|
|
9 |
ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting Transformations
|
|
|
|
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02889823 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States (2020)
|
|
BASE
|
|
Show details
|
|
10 |
Evaluating the reliability of acoustic speech embeddings
|
|
|
|
In: INTERSPEECH 2020 - Annual Conference of the International Speech Communication Association ; https://hal.inria.fr/hal-02977539 ; INTERSPEECH 2020 - Annual Conference of the International Speech Communication Association, Oct 2020, Shanghai / Vitrtual, China (2020)
|
|
BASE
|
|
Show details
|
|
11 |
When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models
|
|
|
|
In: https://hal.inria.fr/hal-03109106 ; 2020 (2020)
|
|
BASE
|
|
Show details
|
|
12 |
Comparing Statistical and Neural Models for Learning Sound Correspondences
|
|
|
|
In: LT4HALA 2020 : First Workshop on Language Technologies for Historical and Ancient Languages ; https://hal.inria.fr/hal-02529929 ; LT4HALA 2020 : First Workshop on Language Technologies for Historical and Ancient Languages, May 2020, Marseille, France (2020)
|
|
BASE
|
|
Show details
|
|
15 |
Can Multilingual Language Models Transfer to an Unseen Dialect? A Case Study on North African Arabizi ...
|
|
|
|
BASE
|
|
Show details
|
|
16 |
Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering ...
|
|
|
|
BASE
|
|
Show details
|
|
17 |
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages ...
|
|
|
|
BASE
|
|
Show details
|
|
18 |
When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models ...
|
|
|
|
BASE
|
|
Show details
|
|
19 |
MUSS: Multilingual Unsupervised Sentence Simplification by Mining Paraphrases ...
|
|
|
|
BASE
|
|
Show details
|
|
20 |
ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting Transformations ...
|
|
|
|
BASE
|
|
Show details
|
|
|
|