DE eng

Search in the Catalogues and Directories

Hits 1 – 8 of 8

1
Universal Segmentations 1.0 (UniSegments 1.0)
Žabokrtský, Zdeněk; Bafna, Nyati; Bodnár, Jan. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2022
BASE
Show details
2
Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0
Kuvač Kraljević, Jelena; Hržica, Gordana; Štefanec, Vanja. - : Jožef Stefan Institute, 2021. : Faculty of Education and Rehabilitation, University of Zagreb, 2021
BASE
Show details
3
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.1
Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
BASE
Show details
4
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.0
Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
BASE
Show details
5
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.0
Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
BASE
Show details
6
The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Serbian 1.0
Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
BASE
Show details
7
The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0
Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
BASE
Show details
8
The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.1
Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
Abstract: The model for lemmatisation of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1183), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~97.54. The difference to the previous version of the lemmatizer is that now it relies solely on XPOS annotations, and not on a combination of UPOS, FEATS (lexicon lookup) and XPOS (lemma prediction) annotations.
Keyword: computer-mediated communication; language model; lemmatisation
URL: http://hdl.handle.net/11356/1352
BASE
Hide details

Catalogues
0
0
0
0
0
0
0
Bibliographies
0
0
0
0
0
0
0
0
0
Linked Open Data catalogues
0
Online resources
0
0
0
0
Open access documents
8
0
0
0
0
© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern