DE eng

Search in the Catalogues and Directories

Page: 1 2
Hits 21 – 35 of 35

21
Universal Dependencies 1.1
Agić, Željko; Aranzabe, Maria Jesus; Atutxa, Aitziber. - : Universal Dependencies Consortium, 2015
BASE
Show details
22
Universal Dependencies 1.2
Nivre, Joakim; Agić, Željko; Aranzabe, Maria Jesus. - : Universal Dependencies Consortium, 2015
BASE
Show details
23
Universal Dependencies 1.0
Nivre, Joakim; Bosco, Cristina; Choi, Jinho. - : Universal Dependencies Consortium, 2015
BASE
Show details
24
Prague Dependency Treebank 3.0
Bejček, Eduard; Hajičová, Eva; Hajič, Jan. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2014
BASE
Show details
25
Prague Czech-English Dependency Treebank 2.0
Abstract: Texts The Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) is a major update of the Prague Czech-English Dependency Treebank 1.0 (LDC2004T25). It is a manually parsed Czech-English parallel corpus sized over 1.2 million running words in almost 50,000 sentences for each part. Data The English part contains the entire Penn Treebank - Wall Street Journal Section (LDC99T42). The Czech part consists of Czech translations of all of the Penn Treebank-WSJ texts. The corpus is 1:1 sentence-aligned. An additional automatic alignment on the node level (different for each annotation layer) is part of this release, too. The original Penn Treebank-like file structure (25 sections, each containing up to one hundred files) has been preserved. Only those PTB documents which have both POS and structural annotation (total of 2312 documents) have been translated to Czech and made part of this release. Each language part is enhanced with a comprehensive manual linguistic annotation in the PDT 2.0 style (LDC2006T01, Prague Dependency Treebank 2.0). The main features of this annotation style are: dependency structure of the content words and coordinating and similar structures (function words are attached as their attribute values) semantic labeling of content words and types of coordinating structures argument structure, including an argument structure ("valency") lexicon for both languages ellipsis and anaphora resolution. This annotation style is called tectogrammatical annotation and it constitutes the tectogrammatical layer in the corpus. For more details see below and documentation. Annotation of the Czech part Sentences of the Czech translation were automatically morphologically annotated and parsed into surface-syntax dependency trees in the PDT 2.0 annotation style. This annotation style is sometimes called analytical annotation; it constitutes the analytical layer of the corpus. The manual tectogrammatical (deep-syntax) annotation was built as a separate layer above the automatic analytical (surface-syntax) parse. A sample of 2,000 sentences was manually annotated on the analytical layer. Annotation of the English part The resulting manual tectogrammatical annotation was built above an automatic transformation of the original phrase-structure annotation of the Penn Treebank into surface dependency (analytical) representations, using the following additional linguistic information from other sources: PropBank (LDC2004T14) VerbNet NomBank (LDC2008T23) flat noun phrase structures (by courtesy of D. Vadas and J.R. Curran) For each sentence, the original Penn Treebank phrase structure trees are preserved in this corpus together with their links to the analytical and tectogrammatical annotation. ; Ministry of Education of the Czech Republic projects No.: MSM0021620838 LC536 ME09008 LM2010013 7E09003+7E11051 7E11041 Czech Science Foundation, grants No.: GAP406/10/0875 GPP406/10/P193 GA405/09/0729 Research funds of the Faculty of Mathematics and Physics, Charles University, Czech Republic, Grant Agency of the Academy of Sciences of the Czech Republic: No. 1ET101120503 Students participating in this project have been running their own student grants from the Grant Agency of the Charles University, which were connected to this project. Only ongoing projects are mentioned: 116310, 158010, 3537/2011 Also, this work was funded in part by the following projects sponsored by the European Commission: Companions, No. 034434 EuroMatrix, No. 034291 EuroMatrixPlus, No. 231720 Faust, No. 247762
Keyword: dependency annotation; parallel corpus; parallel treebank; PCEDT; Penn Treebank; Wall Street Journal; WSJ
URL: http://hdl.handle.net/11858/00-097C-0000-0015-8DAF-4
BASE
Hide details
26
Prague Dependency Treebank 2.5
Bejček, Eduard; Hajič, Jan; Panevová, Jarmila. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2012
BASE
Show details
27
Prague Dependency Treebank 2.0 - sample data
Hajič, Jan; Panevová, Jarmila; Sgall, Petr. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2011
BASE
Show details
28
Prague Dependency Treebank 2.0 (PDT 2.0)
Hajič, Jan; Panevová, Jarmila; Hajičová, Eva. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2011
BASE
Show details
29
CoNLL 2009 Shared Task - Czech Data
Hajič, Jan; Straňák, Pavel; Štěpánek, Jan. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2011
BASE
Show details
30
Arabic computational linguistics
Farghaly, Ali Ahmed Sabry (Hrsg.); Zitouni, Imed; Fraser, Alexander. - Stanford, Calif. : CSLI Publ., 2010
BLLDB
UB Frankfurt Linguistik
Show details
31
Tectogrammatical annotation of the Wall Street Journal
In: The Prague bulletin of mathematical linguistics. - Praha : Univ. (2009) 92, 85-104
BLLDB
OLC Linguistik
Show details
32
The Czech Academic Corpus 2.0 guide
In: The Prague bulletin of mathematical linguistics. - Praha : Univ. (2008) 89, 41-96
BLLDB
OLC Linguistik
Show details
33
Learning to use the Prague Arabic Dependency Treebank
In: Symposium on Arabic Linguistics <19, 2005, Urbana, Ill.>. Perspectives on Arabic linguistics ; 19. Papers from the Nineteenth Annual Symposium on Arabic Linguistics. - Amsterdam [u.a.] : Benjamins (2007), 77-93
BLLDB
Show details
34
The Prague Dependency Treebank : a three-level annotation scenario
In: Treebanks. - Dordrecht [u.a.] : Kluwer (2003), 103-127
BLLDB
Show details
35
Annotation lexicons : using the valency lexicon for tectogrammatical annotation
In: The Prague bulletin of mathematical linguistics. - Praha : Univ. (2003) 79-80, 61-85
BLLDB
Show details

Page: 1 2

Catalogues
1
0
2
0
0
0
0
Bibliographies
6
0
0
0
0
0
0
0
0
Linked Open Data catalogues
0
Online resources
0
0
0
0
Open access documents
29
0
0
0
0
© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern