1 |
The Orange workflow for observing collocation trends ColTrend 1.0
|
|
|
|
BASE
|
|
Show details
|
|
3 |
Slovene ontology of semantic types for nouns SLONEST-noun 1.0
|
|
|
|
BASE
|
|
Show details
|
|
6 |
Multiword Expressions lexicon extracted from the Gigafida 2.1 corpus
|
|
|
|
BASE
|
|
Show details
|
|
7 |
The Orange workflow for observing collocation clusters ColEmbed 1.0
|
|
|
|
BASE
|
|
Show details
|
|
9 |
Frequency lists of collocations from the Gigafida 2.1 corpus
|
|
|
|
BASE
|
|
Show details
|
|
11 |
Language Teachers and Crowdsourcing: Insights from a Cross-European Survey
|
|
|
|
In: ISSN: 1331-6745 ; EISSN: 1849-0379 ; Rasprave Instituta za hrvatski jezik i jezikoslovlje ; https://hal.inria.fr/hal-02974069 ; Rasprave Instituta za hrvatski jezik i jezikoslovlje, 2020, 46 (1), pp.1-28. ⟨10.31724/rihjj.46.1.1⟩ (2020)
|
|
BASE
|
|
Show details
|
|
14 |
Frequency lists of character-level n-grams from the GOS 1.0 corpus 1.1
|
|
|
|
BASE
|
|
Show details
|
|
19 |
Frequency lists of word-level n-grams from the GOS 1.0 corpus 1.1
|
|
|
|
Abstract:
Frequency lists of word-level n-grams (or word sets) were extracted from the GOS 1.0 Corpus of Spoken Slovene (http://hdl.handle.net/11356/1040) using the LIST corpus extraction tool (http://hdl.handle.net/11356/1227). The lists contain all word-level 2-, 3-, 4- and 5-grams occurring in the corpus along with their absolute and relative frequencies, percentages, distribution across the text-types included in the corpus taxonomy, and five collocation measures: Dice, t-score, MI, MI3, logDice, and simple LL. The n-grams were extracted from lower-case word forms, standardized word forms, and morphosyntactic tags. For large lists, shortened versions with the first 150,000 lines were also prepared to facilitate further processing in spreadsheet analysis software. Compared to the previous version (http://hdl.handle.net/11356/1271), this one includes fixes of several typos and substitutes all instances of "normalized forms" with the more adequate term "standardized forms" (as used in the SSJ project).
|
|
Keyword:
morphosyntactic tags; n-grams; Slovenian language; spoken corpus; standardized forms; word forms; word sets; words
|
|
URL: http://hdl.handle.net/11356/1365
|
|
BASE
|
|
Hide details
|
|
|
|