2 |
The CLASSLA-StanfordNLP model for lemmatisation of standard Slovenian 1.4
|
|
|
|
BASE
|
|
Show details
|
|
4 |
The Twitter user dataset for discriminating between Bosnian, Croatian, Montenegrin and Serbian Twitter-HBS 1.0
|
|
|
|
Abstract:
The Twitter-HBS dataset consists of Twitter users, their tweets, and the label of their predominantly used language - Bosnian, Croatian, Montenegrin, or Serbian. Among the tweets, there are also tweets in other languages (mainly English) as the label encodes the predominantly used language of a user only. The main intended usage of this dataset is discrimination between closely-related languages on the level of a Twitter user (not a single tweet). The only pre-processing performed on the texts of the tweets is the transliteration from the Cyrillic into the Latin script so that the dataset measures the quality of the user classifications regardless of the script used.
|
|
Keyword:
closely related languages; language identification; Twitter
|
|
URL: http://hdl.handle.net/11356/1482
|
|
BASE
|
|
Hide details
|
|
8 |
The news dataset for discriminating between Bosnian, Croatian and Serbian SETimes.HBS 1.0
|
|
|
|
BASE
|
|
Show details
|
|
9 |
The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.3
|
|
|
|
BASE
|
|
Show details
|
|
10 |
The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild ...
|
|
|
|
BASE
|
|
Show details
|
|
13 |
Retweet communities reveal the main sources of hate speech
|
|
|
|
In: PLoS One (2022)
|
|
BASE
|
|
Show details
|
|
14 |
The ParlaMint corpora of parliamentary proceedings
|
|
|
|
In: Lang Resour Eval (2022)
|
|
BASE
|
|
Show details
|
|
18 |
Choice of plausible alternatives dataset in Croatian COPA-HR
|
|
|
|
BASE
|
|
Show details
|
|
19 |
Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0
|
|
|
|
BASE
|
|
Show details
|
|
20 |
The Orange workflow for observing collocation trends ColTrend 1.0
|
|
|
|
BASE
|
|
Show details
|
|
|
|