Title |
Domain-related Annotation of Polish Spoken Dialogue Corpus LUNA.PL |
Authors |
Agnieszka Mykowiecka, Katarzyna Głowińska and Joanna Rabiega-Wiśniewska |
Abstract |
The paper presents a corpus of Polish spoken dialogues annotated on several levels, from transcription of dialogues and their morphosyntactic analysis, to semantic annotation. The LUNA.PL corpus is the first semantically annotated corpus of Polish spontaneous speech. It contains 500 dialogues recorded at the Warsaw Transport Authority call centre. For each dialogue, the corpus contains recorded audio signal, its transcription and five XML files with annotations on subsequent levels. Speech transcription was done manually. Text annotation was constructed using a combination of rule based programmes and computer-aided manual work. For morphological annotation we used the already existing analyzer and manually disambiguated the results. Morphologically annotated texts of dialogues were automatically segmented into elementary syntactic chunks. Semantic annotation was done by a set of specially designed rules and then manually corrected. The paper describes details of the domain related semantic annotation which consists of two levels - concept level at which around 200 attributes and their values are annotated, and predicate level at which 47 frame types are recognized. We describe the domain model accepted, and the statistics over the entire annotated set of dialogues. |
Topics |
Dialogue, Speech Recognition/Understanding, Semantics |
Full paper |
Domain-related Annotation of Polish Spoken Dialogue Corpus LUNA.PL |
Slides |
- |
Bibtex |
@InProceedings{MYKOWIECKA10.337,
author = {Agnieszka Mykowiecka and Katarzyna Głowińska and Joanna Rabiega-Wiśniewska}, title = {Domain-related Annotation of Polish Spoken Dialogue Corpus LUNA.PL}, booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)}, year = {2010}, month = {may}, date = {19-21}, address = {Valletta, Malta}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-6-7}, language = {english} } |