Summary of the paper

Title WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words
Authors Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian and Marcello Federico
Abstract This paper presents WAGS (Word Alignment Gold Standard), a novel benchmark which allows extensive evaluation of WA tools on out-of-vocabulary (OOV) and rare words. WAGS is a subset of the Common Test section of the Europarl English-Italian parallel corpus, and is specifically tailored to OOV and rare words. WAGS is composed of 6,715 sentence pairs containing 11,958 occurrences of OOV and rare words up to frequency 15 in the Europarl Training set (5,080 English words and 6,878 Italian words), representing almost 3% of the whole text. Since WAGS is focused on OOV/rare words, manual alignments are provided for these words only, and not for the whole sentences. Two off-the-shelf word aligners have been evaluated on WAGS, and results have been compared to those obtained on an existing benchmark tailored to full text alignment. The results obtained confirm that WAGS is a valuable resource, which allows a statistically sound evaluation of WA systems' performance on OOV and rare words, as well as extensive data analyses. WAGS is publicly released under a Creative Commons Attribution license.
Topics Corpus (Creation, Annotation, etc.), Machine Translation, SpeechToSpeech Translation, Evaluation Methodologies
Full paper WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words
Bibtex @InProceedings{BENTIVOGLI16.607,
  author = {Luisa Bentivogli and Mauro Cettolo and M. Amin Farajian and Marcello Federico},
  title = {WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words},
  booktitle = {Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)},
  year = {2016},
  month = {may},
  date = {23-28},
  location = {Portorož, Slovenia},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Sara Goggi and Marko Grobelnik and Bente Maegaard and Joseph Mariani and Helene Mazo and Asuncion Moreno and Jan Odijk and Stelios Piperidis},
  publisher = {European Language Resources Association (ELRA)},
  address = {Paris, France},
  isbn = {978-2-9517408-9-1},
  language = {english}
 }
Powered by ELDA © 2016 ELDA/ELRA