Title |
A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters |
Authors |
Grzegorz Chrupała and Dietrich Klakow |
Abstract |
Named Entity Recognition is a relatively well-understood NLP task, with many publicly available training resources and software for processing English data. Other languages tend to be underserved in this area. For German, CoNLL-2003 Shared Task provided training data, but there are no publicly available, ready-to-use tools. We fill this gap and develop a German NER system with state-of-the-art performance. In addition to CoNLL 2003 labeled training data, we use two additional resources: (i) 32 million words of unlabeled news article text and (ii) infobox labels from German Wikipedia articles. From the unlabeled text we derive distributional word clusters. Then we use cluster membership features and Wikipedia infobox label features to train a supervised model on the labeled training data. This approach allows us to deal better with word-types unseen in the training data and achieve good performance on German with little engineering effort. |
Topics |
Named Entity recognition, Multilinguality, Tools, systems, applications |
Full paper |
A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters |
Slides |
- |
Bibtex |
@InProceedings{CHRUPAA10.538,
author = {Grzegorz Chrupała and Dietrich Klakow}, title = {A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters}, booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)}, year = {2010}, month = {may}, date = {19-21}, address = {Valletta, Malta}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-6-7}, language = {english} } |