Home Catalogue search

eng

Refine your search:

Search in the Catalogues and Directories






	Sort by
Simple Search

Page: 1 2 3 4 5 6...14

Hits 21 – 40 of 267

21	Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
	Ortiz Suárez, Pedro Javier [Verfasser]; Sagot, Benoît [Verfasser]; Romary, Laurent [Verfasser]. - Mannheim : Leibniz-Institut für Deutsche Sprache (IDS), Bibliothek, 2019
	DNB Subject Category Language
	Show details

22	How OCR Performance can Impact on the Automatic Extraction of Dictionary Content Structures
	Khemakhem, Mohamed; Galleron, Ioana; Williams, Geoffrey...
	In: 19th annual Conference and Members’ Meeting of the Text Encoding Initiative Consortium (TEI) -What is text, really? TEI and beyond ; https://hal.archives-ouvertes.fr/hal-02263276 ; 19th annual Conference and Members’ Meeting of the Text Encoding Initiative Consortium (TEI) -What is text, really? TEI and beyond, Sep 2019, Graz, Austria (2019)
	BASE
	Show details

23	Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
	Ortiz Suárez, Pedro Javier; Sagot, Benoît; Romary, Laurent
	In: 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7) ; https://hal.inria.fr/hal-02148693 ; 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7), Jul 2019, Cardiff, United Kingdom. ⟨10.14618/IDS-PUB-9021⟩ (2019)
	BASE
	Show details

24	Nénufar: Modelling a Diachronic Collection of Dictionary Editions as a Computational Lexical Resource
	Bohbot, Hervé; Frontini, Francesca; Khan, Fahad...
	In: ELEX 2019: smart lexicography ; https://hal.inria.fr/hal-02272978 ; ELEX 2019: smart lexicography, Oct 2019, Sintra, Portugal (2019)
	BASE
	Show details

25	LMF Reloaded
	Romary, Laurent; Khemakhem, Mohamed; Khan, Fahad...
	In: AsiaLex 2019: Past, Present and Future ; https://hal.inria.fr/hal-02118319 ; AsiaLex 2019: Past, Present and Future, Jun 2019, Istanbul, Turkey (2019)
	BASE
	Show details

26	TEI Encoding of a Classical Mixtec Dictionary Using GROBID- Dictionaries
	Bowers, Jack; Khemakhem, Mohamed; Romary, Laurent
	In: ELEX 2019: Smart Lexicography ; https://hal.inria.fr/hal-02264033 ; ELEX 2019: Smart Lexicography, Oct 2019, Sintra, Portugal ; https://elex.link/elex2019/ (2019)
	BASE
	Show details

27	CamemBERT: a Tasty French Language Model
	Martin, Louis; Muller, Benjamin; Ortiz Suárez, Pedro Javier...
	In: https://hal.inria.fr/hal-02445946 ; 2019 (2019)
	BASE
	Show details

28	TEI and the Mixtepec-Mixtec corpus: data integration, annotation and normalization of heterogeneous data for an under-resourced language
	Bowers, Jack; Romary, Laurent
	In: 6th International Conference on Language Documentation and Conservation (ICLDC) ; https://hal.inria.fr/hal-02075475 ; 6th International Conference on Language Documentation and Conservation (ICLDC), Feb 2019, Honolulu, United States (2019)
	BASE
	Show details

29	Preparing the Dictionnaire Universel for Automatic Enrichment
	Ortiz Suárez, Pedro Javier; Romary, Laurent; Sagot, Benoît
	In: 10th International Conference on Historical Lexicography and Lexicology (ICHLL) ; https://hal.inria.fr/hal-02131598 ; 10th International Conference on Historical Lexicography and Lexicology (ICHLL), Jun 2019, Leeuwarden, Netherlands ; https://easychair.org/smart-program/ICHLL-10/ (2019)
	Abstract: International audience ; The Dictionnaire Universel (DU) is an encyclopaedic dictionary originally written by Antoine Furetière around 1676-78, later revised and improved by the Protestant jurist Henri Basnage de Beauval who expanded, corrected and included terms of arts, crafts and sciences, into the Dictionnaire.The aim of the BASNUM project is to digitize the DU in its second edition rewritten by Basnage de Beauval, to analyse it with computational methods in order to better assess the importance of this work for the evolution of sciences and mentalities in the 18th century, and to contribute to the contemporary movement for creating innovative and data-driven computational methods for text digitization, encoding and analysis.Based on the experience acquired within the research group, an enrichment workflow based upon a series of Natural Language Processing processes is being set up to be applied to Basnage's work. This includes, among others, automatic identification of the dictionary structure (macro-, meso- and microstructure), named-entity recognition (in particular persons and locations), classification of dictionary entries, detection and study of polysemy markers, tracking and classification of quotation use (bibliographic references), scoring semantic similarity between the DU and other dictionaries. The main challenges being the lack of available annotated data in order to train machine learning models, decreased accuracy when using modern pre-trained models due to the differences between present-day and 18th century French, and even unreliable or low quality OCRisation. The paper describes methods that are useful to tackle these issues in order to prepare the the DU for automatic enrichment going beyond what current available tools like Grobid-dictionaries can do, thanks to the advent of deep learning NLP models. The paper also describes how these methods could be applied to other dictionaries or even other types of ancient texts.
	Keyword: [INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]; [INFO.INFO-TT]Computer Science [cs]/Document and Text Processing; [SHS.LANGUE]Humanities and Social Sciences/Linguistics; [SHS.MUSEO]Humanities and Social Sciences/Cultural heritage and museology
	URL: https://hal.inria.fr/hal-02131598/document https://hal.inria.fr/hal-02131598 https://hal.inria.fr/hal-02131598/file/ICHLL_10_Slides.pdf
	BASE
	Hide details

30	Connecting the Humanities through Research Infrastructures
	Bassett, Sheena; Wessels, Leon; Krauwer, Steven...
	In: 4th Digital Humanities in the Nordic Countries (DHN 2019) ; https://hal.inria.fr/hal-02047512 ; 4th Digital Humanities in the Nordic Countries (DHN 2019), Mar 2019, Copenhagen, Denmark ; https://cst.dk/DHN2019/DHN2019.html (2019)
	BASE
	Show details

31	The place of lexicography in (computer) science
	Romary, Laurent
	In: The Future of Academic Lexicography: Linguistic Knowledge Codification in the Era of Big Data and AI ; https://hal.inria.fr/hal-02358218 ; The Future of Academic Lexicography: Linguistic Knowledge Codification in the Era of Big Data and AI, Frieda Steurs; Dirk Geeraerts; Niels Schiller; Marian Klamer; Iztok Kosem, Nov 2019, Leiden, Netherlands ; https://www.lorentzcenter.nl/lc/web/2019/1177/program.php3?wsid=1177&venue=Oort (2019)
	BASE
	Show details

32	Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures ...
	Ortiz Suárez, Pedro Javier; Sagot, Benoît; Romary, Laurent. - : Leibniz-Institut für Deutsche Sprache, 2019
	BASE
	Show details

33	From disparate disciplines to unity in diversity. How the PARTHENOS project brings Humanities Research Infrastructures together ...
	Uiterwaal, Frank; Niccolucci, Franco; Bassett, Sheena. - : Zenodo, 2019
	BASE
	Show details

34	From disparate disciplines to unity in diversity. How the PARTHENOS project brings Humanities Research Infrastructures together ...
	Uiterwaal, Frank; Niccolucci, Franco; Bassett, Sheena. - : Zenodo, 2019
	BASE
	Show details

35	LMF Reloaded ...
	Romary, Laurent; Khemakhem, Mohamed; Khan, Fahad. - : arXiv, 2019
	BASE
	Show details

36	LMF Reloaded ...
	Romary, Laurent; Khemakhem, Mohamed; Khan, Anas Fahad. - : Zenodo, 2019
	BASE
	Show details

37	LMF Reloaded ...
	Romary, Laurent; Khemakhem, Mohamed; Khan, Anas Fahad. - : Zenodo, 2019
	BASE
	Show details

38	Automatic TEI encoding of manuscripts catalogues with GROBID-Dictionaries ...
	Noyer, Lucie Rondeau Du; Gabay, Simon; Khmakhem, Mohamed. - : Zenodo, 2019
	BASE
	Show details

39	TEI Lex-0: A Target Format for TEI-Encoded Dictionaries and Lexical Resources ...
	Romary, Laurent; Tasovac, Toma. - : Zenodo, 2019
	BASE
	Show details

40	Automatic TEI encoding of manuscripts catalogues with GROBID-Dictionaries ...
	Noyer, Lucie Rondeau Du; Gabay, Simon; Khmakhem, Mohamed. - : Zenodo, 2019
	BASE
	Show details

Page: 1 2 3 4 5 6...14

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern