Home Catalogue search

eng

Refine your search:
- Keyword:
- Creator / Publisher
- Year:
  - 2013 (7)
- Medium:
  - Online (7)
- Type
- BLLDB-Access:
  - free (7)
  - subject to license (0)

Search in the Catalogues and Directories






	Sort by
Simple Search

Hits 1 – 7 of 7

1	OntoNotes Release 5.0
	Weischedel, Ralph; Palmer, Martha; Marcus, Mitchell. - : Linguistic Data Consortium, 2013. : https://www.ldc.upenn.edu, 2013
	BASE
	Show details

2	Chinese Proposition Bank 3.0
	Xue, Nianwen; Bai, Xiaopeng; Lu, Jill. - : Linguistic Data Consortium, 2013. : https://www.ldc.upenn.edu, 2013
	BASE
	Show details

3	Chinese Treebank 8.0
	Xue, Nianwen; Zhang, Xiuhong; Jiang, Zixin; Palmer, Martha; Xia, Fei; Chiou, Fu-Dong; Chang, Meiyu. - : Linguistic Data Consortium, 2013. : https://www.ldc.upenn.edu, 2013
	Abstract: Introduction Chinese Treebank 8.0 consists of approximately 1.5 million words of annotated and parsed text from Chinese newswire, government documents, magazine articles, various broadcast news and broadcast conversation programs, web newsgroups and weblogs. The Chinese Treebank project began at the University of Pennsylvania in 1998, continued at the University of Colorado and then moved to Brandeis University. The project goal is to provide a large, part-of-speech tagged and fully bracketed Chinese language corpus. The first delivery, Chinese Treebank 1.0, contained 100,000 syntactically annotated words from Xinhua News Agency newswire. It was later corrected and released in 2001 as Chinese Treebank 2.0 (LDC2001T11) and consisted of approximately 100,000 words. LDC released Chinese Treebank 4.0 (LDC2004T05), an updated version containing roughly 400,000 words, in 2004. A year later, LDC published the 500,000 word Chinese Treebank 5.0 (LDC2005T01). Chinese Treebank 6.0 (LDC2007T36), released in 2007, consisted of 780,000 words. Chinese Treebank 7.0 (LDC2010T08), released in 2010, added new annotated newswire data, broadcast material and web text to the approximate total of one million words. Chinese Treebank 8.0 adds new annotated data from newswire, magazine articles and government documents. Data There are 3,007 text files in this release, containing 71,369 sentences, 1,620,561 words, 2,589,848 characters (hanzi or foreign). The data is provided in UTF-8 encoding, and the annotation has Penn Treebank-style labeled brackets. Details of the annotation standard can be found in the segmentation, POS-tagging and bracketing guidelines included in this release. The data is provided in four different formats: raw text, word segmented, POS-tagged and syntactically bracketed formats. All files were automatically verified and manually checked. Samples Please view samples in each format: * POS Tagged * Raw Text * Word Segmented * Syntactically Bracketed Sponsorship This work was supported in part by the Defense Advanced Research Projects Agency GALE Program Grant No. HR0011-06-0022 and BOLT Program No. HR0011-11-C-0145. The content of this publication does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. Updates None at this time.
	URL: https://catalog.ldc.upenn.edu/LDC2013T21
	BASE
	Hide details

4	Chinese Proposition Bank 3.0 ...
	Xue, Nianwen; Bai, Xiaopeng; Lu, Jill. - : Linguistic Data Consortium, 2013
	BASE
	Show details

5	Chinese Treebank 8.0 ...
	Xue, Nianwen; Zhang, Xiuhong; Jiang, Zixin. - : Linguistic Data Consortium, 2013
	BASE
	Show details

6	OntoNotes Release 5.0 ...
	Weischedel, Ralph; Palmer, Martha; Marcus, Mitchell. - : Linguistic Data Consortium, 2013
	BASE
	Show details

7	Chinese Treebank 7.0
	Xue, Nianwen; Jiang, Zixin; Zhong, Xiuhong...
	In: broadcast conversation, broadcast news, news magazine, newswire, web collection (2013)
	BASE
	Show details

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern