LREC 2016 Proceedings

Summary of the paper

Title	Internet Argument Corpus 2.0: An SQL schema for Dialogic Social Media and the Corpora to go with it
Authors	Rob Abbott, Brian Ecker, Pranav Anand and Marilyn Walker
Abstract	Large scale corpora have benefited many areas of research in natural language processing, but until recently, resources for dialogue have lagged behind. Now, with the emergence of large scale social media websites incorporating a threaded dialogue structure, content feedback, and self-annotation (such as stance labeling), there are valuable new corpora available to researchers. In previous work, we released the INTERNET ARGUMENT CORPUS, one of the first larger scale resources available for opinion sharing dialogue. We now release the INTERNET ARGUMENT CORPUS 2.0 (IAC 2.0) in the hope that others will find it as useful as we have. The IAC 2.0 provides more data than IAC 1.0 and organizes it using an extensible, repurposable SQL schema. The database structure in conjunction with the associated code facilitates querying from and combining multiple dialogically structured data sources. The IAC 2.0 schema provides support for forum posts, quotations, markup (bold, italic, etc), and various annotations, including Stanford CoreNLP annotations. We demonstrate the generalizablity of the schema by providing code to import the ConVote corpus.
Topics	Corpus (Creation, Annotation, etc.), Dialogue, LR Infrastructures and Architectures
Full paper	Internet Argument Corpus 2.0: An SQL schema for Dialogic Social Media and the Corpora to go with it
Bibtex	@InProceedings{ABBOTT16.1126, author = {Rob Abbott and Brian Ecker and Pranav Anand and Marilyn Walker}, title = {Internet Argument Corpus 2.0: An SQL schema for Dialogic Social Media and the Corpora to go with it}, booktitle = {Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)}, year = {2016}, month = {may}, date = {23-28}, location = {Portorož, Slovenia}, editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Sara Goggi and Marko Grobelnik and Bente Maegaard and Joseph Mariani and Helene Mazo and Asuncion Moreno and Jan Odijk and Stelios Piperidis}, publisher = {European Language Resources Association (ELRA)}, address = {Paris, France}, isbn = {978-2-9517408-9-1}, language = {english} }