Title |
Collecting and Using Comparable Corpora for Statistical Machine Translation |
Authors |
Inguna Skadiņa, Ahmet Aker, Nikos Mastropavlos, Fangzhong Su, Dan Tufiș, Mateja Verlic, Andrejs Vasiļjevs, Bogdan Babych, Paul Clough, Robert Gaizauskas, Nikos Glaros, Monica Lestari Paramita and Mārcis Pinnis |
Abstract |
Lack of sufficient parallel data for many languages and domains is currently one of the major obstacles to further advancement of automated translation. The ACCURAT project is addressing this issue by researching methods how to improve machine translation systems by using comparable corpora. In this paper we present tools and techniques developed in the ACCURAT project that allow additional data needed for statistical machine translation to be extracted from comparable corpora. We present methods and tools for acquisition of comparable corpora from the Web and other sources, for evaluation of the comparability of collected corpora, for multi-level alignment of comparable corpora and for extraction of lexical and terminological data for machine translation. Finally, we present initial evaluation results on the utility of collected corpora in domain-adapted machine translation and real-life applications. |
Topics |
Corpus (creation, annotation, etc.), Tools, systems, applications, LR national/international projects, organizational/policy issues |
Full paper |
Collecting and Using Comparable Corpora for Statistical Machine Translation |
Bibtex |
@InProceedings{SKADIA12.925,
author = {Inguna Skadiņa and Ahmet Aker and Nikos Mastropavlos and Fangzhong Su and Dan Tufiș and Mateja Verlic and Andrejs Vasiļjevs and Bogdan Babych and Paul Clough and Robert Gaizauskas and Nikos Glaros and Monica Lestari Paramita and Mārcis Pinnis}, title = {Collecting and Using Comparable Corpora for Statistical Machine Translation}, booktitle = {Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)}, year = {2012}, month = {may}, date = {23-25}, address = {Istanbul, Turkey}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Mehmet Uğur Doğan and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, isbn = {978-2-9517408-7-7}, language = {english} } |