LREC 2008 Proceedings

Summary of the paper

Title	System Evaluation on a Named Entity Corpus from Clinical Notes
Authors	Karin Schuler, Vinod Kaggal, James Masanz, Philip Ogren and Guergana Savova
Abstract	This paper presents the evaluation of the dictionary look-up component of Mayo Clinic’s Information Extraction system. The component was tested on a corpus of 160 free-text clinical notes which were manually annotated with the named entity disease. This kind of clinical text presents many language challenges such as fragmented sentences and heavy use of abbreviations and acronyms. The dictionary used for this evaluation was a subset of SNOMED-CT with semantic types corresponding to diseases/disorders without any augmentation. The algorithm achieves an F-score of 0.56 for exact matches and F-scores of 0.76 and 0.62 for right and left-partial matches respectively. Machine learning techniques are currently under investigation to improve this task.
Language	Single language
Topics	Corpus (creation, annotation, etc.), Named Entity recognition, Tools, systems, applications
Full paper	System Evaluation on a Named Entity Corpus from Clinical Notes
Slides	-
Bibtex	@InProceedings{SCHULER08.764, author = {Karin Schuler, Vinod Kaggal, James Masanz, Philip Ogren and Guergana Savova}, title = {System Evaluation on a Named Entity Corpus from Clinical Notes}, booktitle = {Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)}, year = {2008}, month = {may}, date = {28-30}, address = {Marrakech, Morocco}, editor = {Nicoletta Calzolari (Conference Chair), Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-4-0}, note = {http://www.lrec-conf.org/proceedings/lrec2008/}, language = {english} }