Title |
Producing an Encyclopedic Dictionary using Patent Documents |
Authors |
Atsushi Fujii |
Abstract |
Although the World Wide Web has late become an important source to consult for the meaning of words, a number of technical terms related to high technology are not found on the Web. This paper describes a method to produce an encyclopedic dictionary for high-tech terms from patent information. We used a collection of unexamined patent applications published by the Japanese Patent Office as a source corpus. Given this collection, we extracted terms as headword candidates and retrieved applications including those headwords. Then, we extracted paragraph-style descriptions and categorized them into technical domains. We also extracted related terms for each headword. We have produced a dictionary including approximately 400,000 Japanese terms as headwords. We have also implemented an interface with which users can explore our dictionary by reading text descriptions and viewing a related-term graph. |
Language |
|
Topics |
Information Extraction, Information Retrieval, Question Answering, Corpus (creation, annotation, etc.) |
Full paper |
Producing an Encyclopedic Dictionary using Patent Documents |
Slides |
- |
Bibtex |
@InProceedings{FUJII08.519,
author = {Atsushi Fujii},
title = {Producing an Encyclopedic Dictionary using Patent Documents},
booktitle = {Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)},
year = {2008},
month = {may},
date = {28-30},
address = {Marrakech, Morocco},
editor = {Nicoletta Calzolari (Conference Chair), Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Daniel Tapias},
publisher = {European Language Resources Association (ELRA)},
isbn = {2-9517408-4-0},
note = {http://www.lrec-conf.org/proceedings/lrec2008/},
language = {english}
} |