LREC 2008 Proceedings

Summary of the paper

Title	The Construction and Evaluation of Word Space Models
Authors	Yves Peirsman, Simon De Deyne, Kris Heylen and Dirk Geeraerts
Abstract	Semantic similarity is a key issue in many computational tasks. This paper goes into the development and evaluation of two common ways of automatically calculating the semantic similarity between two words. On the one hand, such methods may depend on a manually constructed thesaurus like (Euro)WordNet. Their performance is often evaluated on the basis of a very restricted set of human similarity ratings. On the other hand, corpus-based methods rely on the distribution of two words in a corpus to determine their similarity. Their performance is generally quantified through a comparison with the judgements of the first type of approach. This paper introduces a new Gold Standard of more than 5,000 human intra-category similarity judgements. We show that corpus-based methods often outperform (Euro)WordNet on this data set, and that the use of the latter as a Gold Standard for the former, is thus often far from ideal.
Language	Single language
Topics	Evaluation methodologies, Semantics, Information Extraction, Information Retrieval
Full paper	The Construction and Evaluation of Word Space Models
Slides	The Construction and Evaluation of Word Space Models
Bibtex	@InProceedings{PEIRSMAN08.784, author = {Yves Peirsman, Simon De Deyne, Kris Heylen and Dirk Geeraerts}, title = {The Construction and Evaluation of Word Space Models}, booktitle = {Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)}, year = {2008}, month = {may}, date = {28-30}, address = {Marrakech, Morocco}, editor = {Nicoletta Calzolari (Conference Chair), Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-4-0}, note = {http://www.lrec-conf.org/proceedings/lrec2008/}, language = {english} }