Title |
Challenges in Building a Multilingual Alpine Heritage Corpus |
Authors |
Martin Volk, Noah Bubenhofer, Adrian Althaus, Maya Bangerter, Lenz Furrer and Beni Ruef |
Abstract |
This paper describes our efforts to build a multilingual heritage corpus of alpine texts. Currently we digitize the yearbooks of the Swiss Alpine Club which contain articles in French, German, Italian and Romansch. Articles comprise mountaineering reports from all corners of the earth, but also scientific topics such as topography, geology or glacierology as well as occasional poetry and lyrics. We have already scanned close to 70,000 pages which has resulted in a corpus of 25 million words, 10% of which is a parallel French-German corpus. We have solved a number of challenges in automatic language identification and text structure recognition. Our next goal is to identify the great variety of toponyms (e.g. names of mountains and valleys, glaciers and rivers, trails and cabins) in this corpus, and we sketch how a large gazetteer of Swiss topographical names can be exploited for this purpose. Despite the size of the resource, exact matching leads to a low recall because of spelling variations, language mixtures and partial repetitions. |
Topics |
Corpus (creation, annotation, etc.), Multilinguality, Named Entity recognition |
Full paper |
Challenges in Building a Multilingual Alpine Heritage Corpus |
Slides |
Challenges in Building a Multilingual Alpine Heritage Corpus |
Bibtex |
@InProceedings{VOLK10.110,
author = {Martin Volk and Noah Bubenhofer and Adrian Althaus and Maya Bangerter and Lenz Furrer and Beni Ruef}, title = {Challenges in Building a Multilingual Alpine Heritage Corpus}, booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)}, year = {2010}, month = {may}, date = {19-21}, address = {Valletta, Malta}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-6-7}, language = {english} } |