Publication Date:
2012
abstract:
The paper describes the methodology which is currently being defined for the construction of a "Merged Italian Dependency Treebank" (MIDT) starting from already existing resources. In particular, it reports the results of a case study carried out on two available dependency treebanks, i.e. TUT and ISST-TANL. The issues raised during the comparison of the annotation schemes underlying the two treebanks are discussed and investigated with a particular emphasis on the definition of a set of linguistic categories to be used as a "bridge" between the specific schemes. As an encoding format, the CoNLL de facto standard is used.
Iris type:
04.01 Contributo in Atti di convegno
Keywords:
Syntactic Annotation; Merging of Resources; Dependency Parsing
List of contributors:
Montemagni, Simonetta
Book title:
Proceedings of the LREC 2012 Workshop on Language Resource Merging