Home
Catalogue search
Refine your search:
Keyword
Creator / Publisher:
Ortiz Suárez, Pedro Javier (11)
Romary, Laurent (11)
Sagot, Benoît (10)
Automatic Language Modelling and ANAlysis & Computational Humanities (ALMAnaCH) (9)
Inria de Paris (9)
Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria) (9)
ANR-18-CE38-0003,BASNUM,Numérisation et analyse du Dictionnaire universel de Basnage de Beauval: lexicographie et réseaux scientifiques(2018) (8)
Sorbonne Université (SU) (8)
ANR-19-P3IA-0001,PRAIRIE,PaRis Artificial Intelligence Research InstitutE(2019) (4)
Dupont, Yoann (3)
more
Year
Medium
Type
BLLDB-Access:
free (11)
subject to license (0)
Search in the Catalogues and Directories
All fields
Title
Creator / Publisher
Keyword
Year
AND
OR
AND NOT
All fields
Title
Creator / Publisher
Keyword
Year
AND
OR
AND NOT
All fields
Title
Creator / Publisher
Keyword
Year
AND
OR
AND NOT
All fields
Title
Creator / Publisher
Keyword
Year
AND
OR
AND NOT
All fields
Title
Creator / Publisher
Keyword
Year
Sort by
creator [A → Z]
'
creator [Z → A]
'
publishing year ↑ (asc)
'
publishing year ↓ (desc)
'
title [A → Z]
'
title [Z → A]
'
Simple Search
Hits 1 – 11 of 11
1
Ungoliant: An Optimized Pipeline for the Generation of a Very Large-Scale Multilingual Web Corpus
Abadji, Julien
;
Ortiz Suárez, Pedro Javier
;
Romary, Laurent
...
In: CMLC 2021 - 9th Workshop on Challenges in the Management of Large Corpora ; https://hal.inria.fr/hal-03301590 ; CMLC 2021 - 9th Workshop on Challenges in the Management of Large Corpora, Jul 2021, Limerick / Virtual, Ireland. ⟨10.14618/ids-pub-10468⟩ ; https://www.cl2021.org/ (2021)
BASE
Show details
2
Expanding the content model of annotationBlock
Bartz, Alexandre
;
Janes, Juliette
;
Romary, Laurent
...
In: Next Gen TEI, 2021 - TEI Conference and Members’ Meeting ; https://hal.archives-ouvertes.fr/hal-03380805 ; Next Gen TEI, 2021 - TEI Conference and Members’ Meeting, Oct 2021, Virtual, United States (2021)
BASE
Show details
3
Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus ...
Abadji, Julien
;
Ortiz Suárez, Pedro Javier
;
Romary, Laurent
. - : Leibniz-Institut für Deutsche Sprache, 2021
BASE
Show details
4
Establishing a New State-of-the-Art for French Named Entity Recognition
Ortiz Suárez, Pedro Javier
;
Dupont, Yoann
;
Muller, Benjamin
...
In: LREC 2020 - 12th Language Resources and Evaluation Conference ; https://hal.inria.fr/hal-02617950 ; LREC 2020 - 12th Language Resources and Evaluation Conference, May 2020, Marseille, France ; http://www.lrec-conf.org (2020)
BASE
Show details
5
CamemBERT: a Tasty French Language Model
Martin, Louis
;
Muller, Benjamin
;
Ortiz Suárez, Pedro Javier
...
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02889805 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States. ⟨10.18653/v1/2020.acl-main.645⟩ (2020)
BASE
Show details
6
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
Ortiz Suárez, Pedro Javier
;
Romary, Laurent
;
Sagot, Benoît
In: ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics ; https://hal.inria.fr/hal-02863875 ; ACL 2020 - 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020, Seattle / Virtual, United States. ⟨10.18653/v1/2020.acl-main.156⟩ ; https://acl2020.org (2020)
BASE
Show details
7
How OCR Performance can Impact on the Automatic Extraction of Dictionary Content Structures
Khemakhem, Mohamed
;
Galleron, Ioana
;
Williams, Geoffrey
...
In: 19th annual Conference and Members’ Meeting of the Text Encoding Initiative Consortium (TEI) -What is text, really? TEI and beyond ; https://hal.archives-ouvertes.fr/hal-02263276 ; 19th annual Conference and Members’ Meeting of the Text Encoding Initiative Consortium (TEI) -What is text, really? TEI and beyond, Sep 2019, Graz, Austria (2019)
BASE
Show details
8
Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
Ortiz Suárez, Pedro Javier
;
Sagot, Benoît
;
Romary, Laurent
In: 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7) ; https://hal.inria.fr/hal-02148693 ; 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7), Jul 2019, Cardiff, United Kingdom. ⟨10.14618/IDS-PUB-9021⟩ (2019)
BASE
Show details
9
CamemBERT: a Tasty French Language Model
Martin, Louis
;
Muller, Benjamin
;
Ortiz Suárez, Pedro Javier
...
In: https://hal.inria.fr/hal-02445946 ; 2019 (2019)
BASE
Show details
10
Preparing the Dictionnaire Universel for Automatic Enrichment
Ortiz Suárez, Pedro Javier
;
Romary, Laurent
;
Sagot, Benoît
In: 10th International Conference on Historical Lexicography and Lexicology (ICHLL) ; https://hal.inria.fr/hal-02131598 ; 10th International Conference on Historical Lexicography and Lexicology (ICHLL), Jun 2019, Leeuwarden, Netherlands ; https://easychair.org/smart-program/ICHLL-10/ (2019)
Abstract:
International audience ; The Dictionnaire Universel (DU) is an encyclopaedic dictionary originally written by Antoine Furetière around 1676-78, later revised and improved by the Protestant jurist Henri Basnage de Beauval who expanded, corrected and included terms of arts, crafts and sciences, into the Dictionnaire.The aim of the BASNUM project is to digitize the DU in its second edition rewritten by Basnage de Beauval, to analyse it with computational methods in order to better assess the importance of this work for the evolution of sciences and mentalities in the 18th century, and to contribute to the contemporary movement for creating innovative and data-driven computational methods for text digitization, encoding and analysis.Based on the experience acquired within the research group, an enrichment workflow based upon a series of Natural Language Processing processes is being set up to be applied to Basnage's work. This includes, among others, automatic identification of the dictionary structure (macro-, meso- and microstructure), named-entity recognition (in particular persons and locations), classification of dictionary entries, detection and study of polysemy markers, tracking and classification of quotation use (bibliographic references), scoring semantic similarity between the DU and other dictionaries. The main challenges being the lack of available annotated data in order to train machine learning models, decreased accuracy when using modern pre-trained models due to the differences between present-day and 18th century French, and even unreliable or low quality OCRisation. The paper describes methods that are useful to tackle these issues in order to prepare the the DU for automatic enrichment going beyond what current available tools like Grobid-dictionaries can do, thanks to the advent of deep learning NLP models. The paper also describes how these methods could be applied to other dictionaries or even other types of ancient texts.
Keyword:
[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]
;
[INFO.INFO-TT]Computer Science [cs]/Document and Text Processing
;
[SHS.LANGUE]Humanities and Social Sciences/Linguistics
;
[SHS.MUSEO]Humanities and Social Sciences/Cultural heritage and museology
URL:
https://hal.inria.fr/hal-02131598/document
https://hal.inria.fr/hal-02131598
https://hal.inria.fr/hal-02131598/file/ICHLL_10_Slides.pdf
BASE
Hide details
11
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures ...
Ortiz Suárez, Pedro Javier
;
Sagot, Benoît
;
Romary, Laurent
. - : Leibniz-Institut für Deutsche Sprache, 2019
BASE
Show details
Mobile view
All
Catalogues
UB Frankfurt Linguistik
0
IDS Mannheim
0
OLC Linguistik
0
UB Frankfurt Retrokatalog
0
DNB Subject Category Language
0
Institut für Empirische Sprachwissenschaft
0
Leibniz-Centre General Linguistics (ZAS)
0
Bibliographies
BLLDB
0
BDSL
0
IDS Bibliografie zur deutschen Grammatik
0
IDS Bibliografie zur Gesprächsforschung
0
IDS Konnektoren im Deutschen
0
IDS Präpositionen im Deutschen
0
IDS OBELEX meta
0
MPI-SHH Linguistics Collection
0
MPI for Psycholinguistics
0
Linked Open Data catalogues
Annohub
0
Online resources
Link directory
0
Journal directory
0
Database directory
0
Dictionary directory
0
Open access documents
BASE
11
Linguistik-Repository
0
IDS Publikationsserver
0
Online dissertations
0
Language Description Heritage
0
© 2013 - 2024 Lin|gu|is|tik
|
Imprint
|
Privacy Policy
|
Datenschutzeinstellungen ändern