Étiquetage d'un corpus hétérogène de français médiéval : enjeux et modalités

We have undertaken a morpho-syntactic tagging of the 2.5 millions words of our corpora of medieval texts. The external and internal heterogeneity of the texts make this task a difficult one. As a result, we had to resort to a double strategy. Since there is actually no tool adapted to our corpora, we had first to rely on a programmable tagger in order to categorize a first text. As a second step, and building on the results obtained with the first text, we produced a tagger based on contextal rule learning. Using this latter tool we subsequently tagged a second, quite "similar" (in terms of external criteria) text. This two-step process was then used once again to tag additional texts.The next phase will be to evaluate the heterogeneity of texts according to internal criteria. The correlation of internal and external heterogeneity will enable us to elaborate a "fine-grained" typology of texts.

Data and Resources

Additional Info

Field Value
Source Romance Corpus Linguistics - Corpora and Spoken Language, Tübingen, Gunter Narr Verlag Tübingen
Author Heiden, Serge, Prévost, Sophie
Maintainer CCSD
Last Updated May 9, 2026, 03:42 (UTC)
Created May 9, 2026, 03:42 (UTC)
Identifier halshs-00087995
Language fr
Rights https://about.hal.science/hal-authorisation-v1/
contributor Interactions, Corpus, Apprentissages, Représentations (ICAR) ; École normale supérieure de Lyon (ENS de Lyon) ; Université de Lyon-Université de Lyon-Université Lumière - Lyon 2 (UL2)-Centre National de la Recherche Scientifique (CNRS)
creator Heiden, Serge
date 2002-05-09T00:00:00
harvest_object_id 2530706c-a765-41cc-bcb2-036b624a50cc
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-10-13T00:00:00
set_spec type:COUV