Summary of the paper

Title A Leveled Reading Corpus of Modern Standard Arabic
Authors Muhamed Al Khalil, Hind Saddiki, Nizar Habash and Latifa Alfalasi
Abstract We present a reading corpus in Modern Standard Arabic to enrich the sparse collection of resources that can be leveraged for educational applications. The corpus consists of textbook material from the curriculum of the United Arab Emirates, spanning all 12 grades (1.4 million tokens) and a collection of 129 unabridged works of fiction (5.6 million tokens) all annotated with reading levels from Grade 1 to Post-secondary. We examine reading progression in terms of lexical coverage, and compare the two sub-corpora (curricular, fiction) to others from clearly established genres (news, legal/diplomatic) to measure representation of their respective genres.
Topics Corpus (Creation, Annotation, Etc.), Other, Computer-Assisted Language Learning (Call)
Full paper A Leveled Reading Corpus of Modern Standard Arabic
Bibtex @InProceedings{AL KHALIL18.619,
  author = {Muhamed Al Khalil and Hind Saddiki and Nizar Habash and Latifa Alfalasi},
  title = "{A Leveled Reading Corpus of Modern Standard Arabic}",
  booktitle = {Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)},
  year = {2018},
  month = {May 7-12, 2018},
  address = {Miyazaki, Japan},
  editor = {Nicoletta Calzolari (Conference chair) and Khalid Choukri and Christopher Cieri and Thierry Declerck and Sara Goggi and Koiti Hasida and Hitoshi Isahara and Bente Maegaard and Joseph Mariani and Hélène Mazo and Asuncion Moreno and Jan Odijk and Stelios Piperidis and Takenobu Tokunaga},
  publisher = {European Language Resources Association (ELRA)},
  isbn = {979-10-95546-00-9},
  language = {english}
  }
Powered by ELDA © 2018 ELDA/ELRA