Encoding Syntactic Annotation

There is a need for a general framework for linguistic annotation that is flexible and extensible enough to accommodate different annotation types and different theoretical and practical approaches, while at the same time enabling their representation in a “pivot” format that can serve as the basis for comparative evaluation, merging, and the development of reusable editing and processing tools. To answer this need, we have developed a framework comprised of an abstract model for a variety of different annotation types (e.g., morpho-syntactic tagging, syntactic annotation, co-reference annotation, etc.), which can be instantiated in different ways depending on the annotator's approach and goals. The results have been incorporated into XCES (Ide, et al., 2000a), the XML instantiation of the Corpus Encoding Standard (Ide, 1998a,b), which provides a ready-made, standard encoding format together with a data architecture designed specifically for linguistically annotated corpora.

Data and Resources

Additional Info

Field Value
Source Treebanks: Building and Using Parsed Corpora
Author Ide, Nancy, Romary, Laurent
Maintainer CCSD
Last Updated May 14, 2026, 07:28 (UTC)
Created May 14, 2026, 07:28 (UTC)
Identifier hal-00079163
Language en
Rights https://about.hal.science/hal-authorisation-v1/
contributor Department of Computer Science, Vassar College [NY] ; Vassar College
creator Ide, Nancy
date 2003-05-14T00:00:00
harvest_object_id c4d06be4-eb3b-49cd-97ad-fc7061a6e4d7
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-11-04T00:00:00
set_spec type:COUV