5 |
Parsing with Pretrained Language Models, Multiple Datasets, and Dataset Embeddings ...
|
|
|
|
Abstract:
With an increase of dataset availability, the potential for learning from a variety of data sources has increased. One particular method to improve learning from multiple data sources is to embed the data source during training. This allows the model to learn generalizable features as well as distinguishing features between datasets. However, these dataset embeddings have mostly been used before contextualized transformer-based embeddings were introduced in the field of Natural Language Processing. In this work, we compare two methods to embed datasets in a transformer-based multilingual dependency parser, and perform an extensive evaluation. We show that: 1) embedding the dataset is still beneficial with these models 2) performance increases are highest when embedding the dataset at the encoder level 3) unsurprisingly, we confirm that performance increases are highest for small datasets and datasets with a low baseline score. 4) we show that training on the combination of all datasets performs similarly to ... : Accepted to TLT at SyntaxFest 2021 ...
|
|
Keyword:
Computation and Language cs.CL; FOS Computer and information sciences
|
|
URL: https://arxiv.org/abs/2112.03625 https://dx.doi.org/10.48550/arxiv.2112.03625
|
|
BASE
|
|
Hide details
|
|
6 |
On the Effectiveness of Dataset Embeddings in Mono-lingual,Multi-lingual and Zero-shot Conditions ...
|
|
|
|
BASE
|
|
Show details
|
|
7 |
Genre as Weak Supervision for Cross-lingual Dependency Parsing ...
|
|
|
|
BASE
|
|
Show details
|
|
9 |
Genre as Weak Supervision for Cross-lingual Dependency Parsing ...
|
|
|
|
BASE
|
|
Show details
|
|
10 |
DaN+: Danish Nested Named Entities and Lexical Normalization ...
|
|
|
|
BASE
|
|
Show details
|
|
11 |
From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding ...
|
|
|
|
BASE
|
|
Show details
|
|
12 |
From Masked-Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding ...
|
|
|
|
BASE
|
|
Show details
|
|
13 |
Lexical Normalization for Code-switched Data and its Effect on POS-tagging ...
|
|
|
|
BASE
|
|
Show details
|
|
14 |
Fair Is Better than Sensational: Man Is to Doctor as Woman Is to Doctor
|
|
|
|
In: Computational Linguistics, Vol 46, Iss 2, Pp 487-497 (2020) (2020)
|
|
BASE
|
|
Show details
|
|
15 |
Bleaching Text: Abstract Features for Cross-lingual Gender Prediction ...
|
|
|
|
BASE
|
|
Show details
|
|
|
|