Home Catalogue search

eng

Refine your search:

Search in the Catalogues and Directories






	Sort by
Simple Search

Page: 1 2 3 4 5

Hits 1 – 20 of 81

1	Abstracts from the KAS corpus KAS-Abs 2.0
	Žagar, Aleš; Kavaš, Matic; Robnik-Šikonja, Marko. - : Faculty of Electrical Engineering and Computer Science, University of Maribor, 2022. : Faculty of Computer and Information Science, University of Ljubljana, 2022
	BASE
	Show details

2	Corpus of academic Slovene KAS 2.0
	Žagar, Aleš; Kavaš, Matic; Robnik-Šikonja, Marko. - : Faculty of Electrical Engineering and Computer Science, University of Maribor, 2022. : Faculty of Computer and Information Science, University of Ljubljana, 2022
	BASE
	Show details

3	Summarization datasets from the KAS corpus KAS-Sum 1.0
	Žagar, Aleš; Kavaš, Matic; Robnik-Šikonja, Marko. - : Faculty of Electrical Engineering and Computer Science, University of Maribor, 2022. : Faculty of Computer and Information Science, University of Ljubljana, 2022
	BASE
	Show details

4	Machine Translation datasets from the KAS corpus KAS-MT 1.0
	Žagar, Aleš; Kavaš, Matic; Robnik-Šikonja, Marko. - : Faculty of Electrical Engineering and Computer Science, University of Maribor, 2022. : Faculty of Computer and Information Science, University of Ljubljana, 2022
	BASE
	Show details

5	The ParlaMint corpora of parliamentary proceedings
	Erjavec, Tomaž; Ogrodniczuk, Maciej; Osenova, Petya. - 2022
	BASE
	Show details

6	The ParlaMint corpora of parliamentary proceedings
	Erjavec, Tomaž; Ogrodniczuk, Maciej; Osenova, Petya...
	In: Lang Resour Eval (2022)
	BASE
	Show details

7	Universal Dependencies 2.9
	Zeman, Daniel; Nivre, Joakim; Abrams, Mitchell. - : Universal Dependencies Consortium, 2021
	BASE
	Show details

8	Universal Dependencies 2.8.1
	Zeman, Daniel; Nivre, Joakim; Abrams, Mitchell. - : Universal Dependencies Consortium, 2021
	BASE
	Show details

9	Universal Dependencies 2.8
	Zeman, Daniel; Nivre, Joakim; Abrams, Mitchell. - : Universal Dependencies Consortium, 2021
	BASE
	Show details

10	Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.0
	Ljubešić, Nikola; Fišer, Darja; Erjavec, Tomaž. - : Jožef Stefan Institute, 2021
	BASE
	Show details

11	Montenegrin web corpus meWaC 1.0
	Ljubešić, Nikola; Erjavec, Tomaž. - : Jožef Stefan Institute, 2021
	BASE
	Show details

12	Comparable corpora of South-Slavic Wikipedias CLASSLA-Wikipedia 1.0
	Ljubešić, Nikola; Markoski, Filip; Markoska, Elena. - : Jožef Stefan Institute, 2021
	BASE
	Show details

13	Multilingual comparable corpora of parliamentary debates ParlaMint 2.1
	Erjavec, Tomaž; Ogrodniczuk, Maciej; Osenova, Petya. - : CLARIN ERIC, 2021
	BASE
	Show details

14	Corpus of Croatian news portals ENGRI (2014-2018)
	Bogunović, Irena; Kučić, Mario; Ljubešić, Nikola. - : University of Rijeka, Faculty of Maritime Studies, 2021
	BASE
	Show details

15	Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.1
	Ljubešić, Nikola; Fišer, Darja; Erjavec, Tomaž; Šulc, Ajda. - : Jožef Stefan Institute, 2021
	Abstract: The FRENK dataset consists of comments to Facebook posts (news articles) of mainstream media outlets from Croatia, Great Britain, and Slovenia, on the topics of migrants and LGBT. The dataset contains whole discussion threads. Each comment is annotated by the type of socially unacceptable discourse (e.g., inappropriate, offensive, violent speech) and its target (e.g., migrants/LGBT, commenters, media). The annotation schema in its details is described in https://arxiv.org/pdf/1906.02045.pdf. Usernames in the metadata are pseudo-anonymised and removed from the comments. The data in each language (Croatian (hr), English (en), Slovenian (sl), and topic (migrants, LGBT) is divided into a training and a testing portion. The training and testing data consist of separate discussion threads, i.e., there is no cross-discussion-thread contamination between training and testing data. The sizes of the splits are the following: Croatian, migrants: 4356 training comments, 978 testing comments; Croatian LGBT: 4494 training comments, 1142 comments; English, migrants: 4540 training comments, 1285 testing comments; English, LGBT: 4819 training comments, 1017 testing comments; Slovenian, migrants: 5145 training comments, 1277 testing comments; Slovenian, LGBT: 2842 training comments, 900 testing comments. The difference to the first version of the dataset are the additions of 1. the annotation guidelines in English and 2. the link to the huggingface dataset.
	Keyword: hate speech; news comments; offensive language
	URL: http://hdl.handle.net/11356/1462
	BASE
	Hide details

16	Abstracts from the KAS corpus KAS-Abs 1.0
	Erjavec, Tomaž; Fišer, Darja; Ljubešić, Nikola. - : Jožef Stefan Institute, 2021. : Faculty of Electrical Engineering and Computer Science, University of Maribor, 2021
	BASE
	Show details

17	Linguistically annotated multilingual comparable corpora of parliamentary debates ParlaMint.ana 2.1
	Erjavec, Tomaž; Ogrodniczuk, Maciej; Osenova, Petya. - : CLARIN ERIC, 2021
	BASE
	Show details

18	Linguistically annotated multilingual comparable corpora of parliamentary debates ParlaMint.ana 2.0
	Erjavec, Tomaž; Ogrodniczuk, Maciej; Osenova, Petya. - : CLARIN ERIC, 2021
	BASE
	Show details

19	Multilingual comparable corpora of parliamentary debates ParlaMint 2.0
	Erjavec, Tomaž; Ogrodniczuk, Maciej; Osenova, Petya. - : CLARIN ERIC, 2021
	BASE
	Show details

20	Corpus of Written Standard Slovene Gigafida 2.0
	Krek, Simon; Erjavec, Tomaž; Repar, Andraž. - : Centre for Language Resources and Technologies, University of Ljubljana, 2021
	BASE
	Show details

Page: 1 2 3 4 5

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern