Title |
Korp ― the corpus infrastructure of Spräkbanken |
Authors |
Lars Borin, Markus Forsberg and Johan Roxendal |
Abstract |
We present Korp, the corpus infrastructure of Spräkbanken (the Swedish Language Bank). The infrastructure consists of three main components: the Korp corpus pipeline, the Korp backend, and the Korp frontend. The Korp corpus pipeline is used for importing corpora, annotating them, and then exporting the annotated corpora into different formats. An essential feature of the pipeline is the ability to leave existing annotations untouched, both structural and word level annotations, and to use the existing annotations as the foundation of other annotations. The Korp backend consists of a set of REST-based web services for searching in and retrieving information about the corpora. Finally, the Korp frontend is a graphical search interface that interacts with the Korp backend. The interface has been inspired by corpus search interfaces such as SketchEngine, Glossa, and DeepDict, and it uses State Chart XML (SCXML) in order to enable users to bookmark interaction states. We give a functional and technical overview of the three components, followed by a discussion of planned future work. |
Topics |
Corpus (creation, annotation, etc.), LR Infrastructures and Architectures, Web Services |
Full paper |
Korp ― the corpus infrastructure of Spräkbanken |
Bibtex |
@InProceedings{BORIN12.248,
author = {Lars Borin and Markus Forsberg and Johan Roxendal}, title = {Korp ― the corpus infrastructure of Spräkbanken}, booktitle = {Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)}, year = {2012}, month = {may}, date = {23-25}, address = {Istanbul, Turkey}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Mehmet Uğur Doğan and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, isbn = {978-2-9517408-7-7}, language = {english} } |