Title |
Transcription Methods for Consistency, Volume and Efficiency |
Authors |
Meghan Lammie Glenn, Stephanie M. Strassel, Haejoong Lee, Kazuaki Maeda, Ramez Zakhary and Xuansong Li |
Abstract |
This paper describes recent efforts at Linguistic Data Consortium at the University of Pennsylvania to create manual transcripts as a shared resource for human language technology research and evaluation. Speech recognition and related technologies in particular call for substantial volumes of transcribed speech for use in system development, and for human gold standard references for evaluating performance over time. Over the past several years LDC has developed a number of transcription approaches to support the varied goals of speech technology evaluation programs in multiple languages and genres. We describe each transcription method in detail, and report on the results of a comparative analysis of transcriber consistency and efficiency, for two transcription methods in three languages and five genres. Our findings suggest that transcripts for planned speech are generally more consistent than those for spontaneous speech, and that careful transcription methods result in higher rates of agreement when compared to quick transcription methods. We conclude with a general discussion of factors contributing to transcription quality, efficiency and consistency. |
Topics |
Corpus (creation, annotation, etc.) |
Full paper |
Transcription Methods for Consistency, Volume and Efficiency |
Slides |
Transcription Methods for Consistency, Volume and Efficiency |
Bibtex |
@InProceedings{GLENN10.849,
author = {Meghan Lammie Glenn and Stephanie M. Strassel and Haejoong Lee and Kazuaki Maeda and Ramez Zakhary and Xuansong Li}, title = {Transcription Methods for Consistency, Volume and Efficiency}, booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)}, year = {2010}, month = {may}, date = {19-21}, address = {Valletta, Malta}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-6-7}, language = {english} } |