Summary of the paper

Title OAL: A NLP Architecture to Improve the Development of Linguistic Resources for NLP
Authors Javier Couto, Helena Blancafort, Somara Seng, Nicolas Kuchmann-Beauger, Anass Talby and Claude de Loupy
Abstract The performance of most NLP applications relies upon the quality of linguistic resources. The creation, maintenance and enrichment of those resources are a labour-intensive task, especially when no tools are available. In this paper we present the NLP architecture OAL, designed to assist computational linguists in the whole process of the development of resources in an industrial context: from corpora compilation to quality assurance. To add new words more easily to the morphosyntactic lexica, a guesser that lemmatizes and assigns morphosyntactic tags as well as inflection paradigms to a new word has been developed. Moreover, different control mechanisms are set up to check the coherence and consistency of the resources. Today OAL manages resources in five European languages: French, English, Spanish, Italian and Polish. Chinese and Portuguese are in process. The development of OAL has followed an incremental strategy. At present, semantic lexica, a named entities guesser and a named entities phonetizer are being developed.
Topics Lexicon, lexical database, Tools, systems, applications, LR Infrastructures and Architectures
Full paper OAL: A NLP Architecture to Improve the Development of Linguistic Resources for NLP
Slides -
Bibtex @InProceedings{COUTO10.882,
  author = {Javier Couto and Helena Blancafort and Somara Seng and Nicolas Kuchmann-Beauger and Anass Talby and Claude de Loupy},
  title = {OAL: A NLP Architecture to Improve the Development of Linguistic Resources for NLP},
  booktitle = {Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)},
  year = {2010},
  month = {may},
  date = {19-21},
  address = {Valletta, Malta},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Bente Maegaard and Joseph Mariani and Jan Odijk and Stelios Piperidis and Mike Rosner and Daniel Tapias},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {2-9517408-6-7},
  language = {english}
 }
Powered by ELDA © 2010 ELDA/ELRA