LREC 2008 Proceedings

Summary of the paper

Title	An Inverted Index for Storing and Retrieving Grammatical Dependencies
Authors	Michaela Atterer and Hinrich Schütze
Abstract	Web count statistics gathered from search engines have been widely used as a resource in a variety of NLP tasks. For some tasks, however, the information they exploit is not fine-grained enough. We propose an inverted index over grammatical relations as a fast and reliable resource to access more general and also more detailed frequency information. To build the index, we use a dependency parser to parse a large corpus. We extract binary dependency relations, such as he-subj-say (“he” is the subject of “say”) as index terms and construct the index using publicly available open-source indexing software. The unit we index over is the sentence. The index can be used to extract grammatical relations and frequency counts for these relations. The framework also provides the possibility to search for partial dependencies (say, the frequency of “he” occurring in subject position), words, strings and a combination of these. One possible application is the disambiguation of syntactic structures.
Language	Language-independent
Topics	Information Extraction, Information Retrieval, LR Infrastructures and Architectures
Full paper	An Inverted Index for Storing and Retrieving Grammatical Dependencies
Slides	-
Bibtex	@InProceedings{ATTERER08.23, author = {Michaela Atterer and Hinrich Schütze}, title = {An Inverted Index for Storing and Retrieving Grammatical Dependencies}, booktitle = {Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)}, year = {2008}, month = {may}, date = {28-30}, address = {Marrakech, Morocco}, editor = {Nicoletta Calzolari (Conference Chair), Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Daniel Tapias}, publisher = {European Language Resources Association (ELRA)}, isbn = {2-9517408-4-0}, note = {http://www.lrec-conf.org/proceedings/lrec2008/}, language = {english} }