Title |
Development and Application of a Cross-language Document Comparability Metric |
Authors |
Fangzhong Su and Bogdan Babych |
Abstract |
In this paper we present a metric that measures comparability of documents across different languages. The metric is developed within the FP7 ICT ACCURAT project, as a tool for aligning comparable corpora on the document level; further these aligned comparable documents are used for phrase alignment and extraction of translation equivalents, with the aim to extend phrase tables of statistical MT systems without the need to use parallel texts. The metric uses several features, such as lexical information, document structure, keywords and named entities, which are combined in an ensemble manner. We present the results by measuring the reliability and effectiveness of the metric, and demonstrate its application and the impact for the task of parallel phrase extraction from comparable corpora. |
Topics |
Machine Translation, SpeechToSpeech Translation, Corpus (creation, annotation, etc.), Evaluation methodologies |
Full paper |
Development and Application of a Cross-language Document Comparability Metric |
Bibtex |
@InProceedings{SU12.804,
author = {Fangzhong Su and Bogdan Babych}, title = {Development and Application of a Cross-language Document Comparability Metric}, booktitle = {Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)}, year = {2012}, month = {may}, date = {23-25}, address = {Istanbul, Turkey}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Mehmet Uğur Doğan and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, isbn = {978-2-9517408-7-7}, language = {english} } |