Home Catalogue search

eng

Refine your search:

Search in the Catalogues and Directories






	Sort by
Simple Search

Hits 1 – 9 of 9

1	IndoNLI: A Natural Language Inference Dataset for Indonesian ...
	Mahendra, Rahmad; Aji, Alham Fikri; Louvan, Samuel. - : arXiv, 2021
	BASE
	Show details

2	IndoNLI: A Natural Language Inference Dataset for Indonesian ...
	The 2021 Conference on Empirical Methods in Natural Language Processing 2021; Aji, Alham Fikri; Louvan, Samuel. - : Underline Science Inc., 2021
	BASE
	Show details

3	What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? ...
	The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021; ., Harsh; Bowman, Samuel R.. - : Underline Science Inc., 2021
	BASE
	Show details

4	Comparing Test Sets with Item Response Theory ...
	The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021; Bowman, Samuel R.; Cho, Kyunghyun. - : Underline Science Inc., 2021
	BASE
	Show details

5	VisualSem: A High-quality Knowledge Graph for Vision and Language ...
	Alberts, Houda; Huang, Teresa; Deshpande, Yash. - : arXiv, 2020
	BASE
	Show details

6	On understanding character-level models for representing morphology ...
	Vania, Clara. - : The University of Edinburgh, 2020
	BASE
	Show details

7	On understanding character-level models for representing morphology
	Vania, Clara. - : The University of Edinburgh, 2020
	BASE
	Show details

8	LINSPECTOR: Multilingual Probing Tasks for Word Representations
	Şahin, Gözde Gül; Vania, Clara; Kuznetsov, Ilia; Gurevych, Iryna
	In: Computational Linguistics, Vol 46, Iss 2, Pp 335-385 (2020) (2020)
	Abstract: Despite an ever-growing number of word representation models introduced for a large number of languages, there is a lack of a standardized technique to provide insights into what is captured by these models. Such insights would help the community to get an estimate of the downstream task performance, as well as to design more informed neural architectures, while avoiding extensive experimentation that requires substantial computational resources not all researchers have access to. A recent development in NLP is to use simple classification tasks, also called probing tasks, that test for a single linguistic feature such as part-of-speech. Existing studies mostly focus on exploring the linguistic information encoded by the continuous representations of English text. However, from a typological perspective the morphologically poor English is rather an outlier: The information encoded by the word order and function words in English is often stored on a subword, morphological level in other languages. To address this, we introduce 15 type-level probing tasks such as case marking, possession, word length, morphological tag count, and pseudoword identification for 24 languages. We present a reusable methodology for creation and evaluation of such tests in a multilingual setting, which is challenging because of a lack of resources, lower quality of tools, and differences among languages. We then present experiments on several diverse multilingual word embedding models, in which we relate the probing task performance for a diverse set of languages to a range of five classic NLP tasks: POS-tagging, dependency parsing, semantic role labeling, named entity recognition, and natural language inference. We find that a number of probing tests have significantly high positive correlation to the downstream tasks, especially for morphologically rich languages. We show that our tests can be used to explore word embeddings or black-box neural models for linguistic cues in a multilingual setting. We release the probing data sets and the evaluation suite LINSPECTOR with https://github.com/UKPLab/linspector .
	Keyword: Computational linguistics. Natural language processing; P98-98.5
	URL: https://doi.org/10.1162/coli_a_00376 https://doaj.org/article/59f9f3ff91d3454eb0a319a7567c3d11
	BASE
	Hide details

9	CoNLL 2017 Shared Task System Outputs
	Zeman, Daniel; Potthast, Martin; Straka, Milan. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2017
	BASE
	Show details

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern