LREC 2010 Proceedings

Summary of the paper

Title	A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters
Authors	Grzegorz Chrupała and Dietrich Klakow
Abstract	Named Entity Recognition is a relatively well-understood NLP task, with many publicly available training resources and software for processing English data. Other languages tend to be underserved in this area. For German, CoNLL-2003 Shared Task provided training data, but there are no publicly available, ready-to-use tools. We fill this gap and develop a German NER system with state-of-the-art performance. In addition to CoNLL 2003 labeled training data, we use two additional resources: (i) 32 million words of unlabeled news article text and (ii) infobox labels from German Wikipedia articles. From the unlabeled text we derive distributional word clusters. Then we use cluster membership features and Wikipedia infobox label features to train a supervised model on the labeled training data. This approach allows us to deal better with word-types unseen in the training data and achieve good performance on German with little engineering effort.
Topics	Named Entity recognition, Multilinguality, Tools, systems, applications
Full paper	A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters
Slides	-
Bibtex	@InProceedings{CHRUPAA10.538, author = {Grzegorz Chrupała and Dietrich Klakow}, title = {A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters}, booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)}, year = {2010}, month = {may}, date = {19-21}, address = {Valletta, Malta}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-6-7}, language = {english} }