Title |
The Romanian Neuter Examined Through A Two-Gender N-Gram Classification System |
Authors |
Liviu P. Dinu, Vlad Niculae and Octavia-Maria Şulea |
Abstract |
Romanian has been traditionally seen as bearing three lexical genders: masculine, feminine and neuter, although it has always been known to have only two agreement patterns (for masculine and feminine). A recent analysis of the Romanian gender system described in (Bateman and Polinsky, 2010), based on older observations, argues that there are two lexically unspecified noun classes in the singular and two different ones in the plural and that what is generally called neuter in Romanian shares the class in the singular with masculines, and the class in the plural with feminines based not only on agreement features but also on form. Previous machine learning classifiers that have attempted to discriminate Romanian nouns according to gender have so far taken as input only the singular form, presupposing the traditional tripartite analysis. We propose a classifier based on two parallel support vector machines using n-gram features from the singular and from the plural which outperforms previous classifiers in its high ability to distinguish the neuter. The performance of our system suggests that the two-gender analysis of Romanian, on which it is based, is on the right track. |
Topics |
Statistical and machine learning methods, Language modelling, Morphology |
Full paper |
The Romanian Neuter Examined Through A Two-Gender N-Gram Classification System |
Bibtex |
@InProceedings{DINU12.651,
author = {Liviu P. Dinu and Vlad Niculae and Octavia-Maria Şulea}, title = {The Romanian Neuter Examined Through A Two-Gender N-Gram Classification System}, booktitle = {Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)}, year = {2012}, month = {may}, date = {23-25}, address = {Istanbul, Turkey}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Mehmet Uğur Doğan and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, isbn = {978-2-9517408-7-7}, language = {english} } |