DE eng

Search in the Catalogues and Directories

Hits 1 – 8 of 8

1
Word-embedding based bilingual terminology alignment ...
BASE
Show details
2
Aligning Estonian and Russian news industry keywords with the help of subtitle translations and an environmental thesaurus ...
Repar, Andraž; Shumakov, Andrej. - : Zenodo, 2021
BASE
Show details
3
Word-embedding based bilingual terminology alignment ...
Abstract: The ability to accurately align concepts between languages can provide significant benefits in many practical applications. In this paper, we extend a machine learning approach using dictionary and cognate-based features with novel cross-lingual embedding features using pretrained fastText embeddings. We use the tool VecMap to align the embeddings between Slovenian and English and then for every word calculate the top 3 closest word embeddings in the opposite language based on cosine distance. These alignments are then used as features for the machine learning algorithm. With one configuration of the input parameters, we managed to improve the overall F-score compared to previous work, while another configuration yielded improved precision (96%) at a cost of lower recall. Using embedding-based features as a replacement for dictionary-based features provides a significant benefit: while a large bilingual parallel corpus is required to generate the Giza++ word alignment lists, no such data is required for ...
Keyword: embeddings alignment; machine learning; terminology alignment; word embeddings
URL: https://zenodo.org/record/5547982
https://dx.doi.org/10.5281/zenodo.5547982
BASE
Hide details
4
Aligning Estonian and Russian news industry keywords with the help of subtitle translations and an environmental thesaurus ...
Repar, Andraž; Shumakov, Andrej. - : Zenodo, 2021
BASE
Show details
5
Corpus of Written Standard Slovene Gigafida 2.0
Krek, Simon; Erjavec, Tomaž; Repar, Andraž. - : Centre for Language Resources and Technologies, University of Ljubljana, 2021
BASE
Show details
6
SimLex-999 Slovenian translation SimLex-999-sl 1.0
Pollak, Senja; Vulić, Ivan; Pelicon, Andraž. - : University of Ljubljana, 2021
BASE
Show details
7
Evaluation of contextual embeddings on less-resourced languages ...
BASE
Show details
8
Reproduction, replication, analysis and adaptation of a term alignment approach [<Journal>]
Repar, Andraž [Verfasser]; Martinc, Matej [Verfasser]; Pollak, Senja [Verfasser]
DNB Subject Category Language
Show details

Catalogues
0
0
0
0
1
0
0
Bibliographies
0
0
0
0
0
0
0
0
0
Linked Open Data catalogues
0
Online resources
0
0
0
0
Open access documents
7
0
0
0
0
© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern