1 |
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.1
|
|
|
|
BASE
|
|
Show details
|
|
2 |
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.0
|
|
|
|
BASE
|
|
Show details
|
|
3 |
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.0
|
|
|
|
BASE
|
|
Show details
|
|
4 |
The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Serbian 1.0
|
|
|
|
BASE
|
|
Show details
|
|
5 |
The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0
|
|
|
|
Abstract:
This model for morphosyntactic annotation of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1210), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1205). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~95.11.
|
|
Keyword:
computer-mediated communication; language model; part-of-speech tagging
|
|
URL: http://hdl.handle.net/11356/1331
|
|
BASE
|
|
Hide details
|
|
6 |
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.1
|
|
|
|
BASE
|
|
Show details
|
|
|
|