Title |
Specifying Treebanks, Outsourcing Parsebanks: FinnTreeBank 3 |
Authors |
Atro Voutilainen, Kristiina Muhonen, Tanja Purtonen and Krister Lindén |
Abstract |
Corpus-based treebank annotation is known to result in incomplete coverage of mid- and low-frequency linguistic constructions: the linguistic representation and corpus annotation quality are sometimes suboptimal. Large descriptive grammars cover also many mid- and low-frequency constructions. We argue for use of large descriptive grammars and their sample sentences as a basis for specifying higher-coverage grammatical representations. We present an sample case from an ongoing project (FIN-CLARIN FinnTreeBank) where an grammatical representation is documented as an annotator's manual alongside manual annotation of sample sentences extracted from a large descriptive grammar of Finnish. We outline the linguistic representation (morphology and dependency syntax) for Finnish, and show how the resulting `Grammar Definition Corpus' and the documentation is used as a task specification for an external subcontractor for building a parser engine for use in morphological and dependency syntactic analysis of large volumes of Finnish for parsebanking purposes. The resulting corpus, FinnTreeBank 3, is due for release in June 2012, and will contain tens of millions of words from publicly available corpora of Finnish with automatic morphological and dependency syntactic analysis, for use in research on the corpus linguistics and language engineering. |
Topics |
Corpus (creation, annotation, etc.), Language modelling, Tools, systems, applications |
Full paper |
Specifying Treebanks, Outsourcing Parsebanks: FinnTreeBank 3 |
Bibtex |
@InProceedings{VOUTILAINEN12.766,
author = {Atro Voutilainen and Kristiina Muhonen and Tanja Purtonen and Krister Lindén}, title = {Specifying Treebanks, Outsourcing Parsebanks: FinnTreeBank 3}, booktitle = {Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)}, year = {2012}, month = {may}, date = {23-25}, address = {Istanbul, Turkey}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Mehmet Uğur Doğan and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, isbn = {978-2-9517408-7-7}, language = {english} } |