Title |
A Corpus for Evaluating Semantic Multilingual Web Retrieval Systems: The Sense Folder Corpus |
Authors |
Ernesto William De Luca |
Abstract |
In this paper, we present the multilingual Sense Folder Corpus. After the analysis of different corpora, we describe the requirements that have to be satisfied for evaluating semantic multilingual retrieval approaches. Justified by the unfulfilled requirements explained, we start creating a small bilingual hand-tagged corpus of 502 documents retrieved from Web searches. The documents contained in this collection have been created using Google queries. A single ambiguous word has been searched and related documents (approx. the first 60 documents for every keyword) have been retrieved. The document collection has been extended at the query word level, using single ambiguous words for English (argument, bank, chair, network and rule) and for Italian (argomento, lingua, regola, rete and stampa). The search and annotation process has been done both in a monolingual way for the English and the Italian language. 252 English and 250 Italian documents have been retrieved from Google and saved in their original rank. The performance of semantic multilingual retrieval systems has been evaluated using such a corpus with three baselines (Random, First Sense and Most Frequent Sense) that are formally presented and discussed. The fine-grained evaluation of the Sense Folder approach is discussed in details. |
Topics |
Corpus (creation, annotation, etc.), Document Classification, Text categorisation, Information Extraction, Information Retrieval |
Full paper |
A Corpus for Evaluating Semantic Multilingual Web Retrieval Systems: The Sense Folder Corpus |
Slides |
- |
Bibtex |
@InProceedings{DELUCA10.816,
author = {Ernesto William De Luca}, title = {A Corpus for Evaluating Semantic Multilingual Web Retrieval Systems: The Sense Folder Corpus}, booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)}, year = {2010}, month = {may}, date = {19-21}, address = {Valletta, Malta}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-6-7}, language = {english} } |