Summary of the paper

Title Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information
Authors Chi-Hsin Yu, Yi-jie Tang and Hsin-Hsi Chen
Abstract Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variations in the web texts, such as character encoding in processing Chinese web texts. In this paper, we aim to develop a web-scale Chinese word N-gram corpus with parts of speech information called NTU PN-Gram corpus using the ClueWeb09 dataset. We focus on the character encoding and some Chinese-specific issues. The statistics about the dataset is reported. We will make the resulting corpus a public available resource to boost the Chinese language processing.
Topics Corpus (creation, annotation, etc.), Language Identification, Tools, systems, applications
Full paper Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information
Bibtex @InProceedings{YU12.596,
  author = {Chi-Hsin Yu and Yi-jie Tang and Hsin-Hsi Chen},
  title = {Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information},
  booktitle = {Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)},
  year = {2012},
  month = {may},
  date = {23-25},
  address = {Istanbul, Turkey},
  editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Mehmet Uğur Doğan and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {978-2-9517408-7-7},
  language = {english}
 }
Powered by ELDA © 2012 ELDA/ELRA