Title

CST Bank: A Corpus for the Study of Cross-document Structural Relationships

Author(s)

Dragomir Radev, Jahna Otterbacher, Zhu Zhang

University of Michigan

Session

P19-SW

Abstract

Clusters of multiple news stories related to the same topic exhibit a number of interesting properties. For example, when documents have been published at various points in time or by different authors or news agencies, one finds many instances of paraphrasing, information overlap and even contradiction. The current paper presents the Cross-document Structure Theory (CST) Bank, a collection of multi-document clusters in which pairs of sentences from different documents have been annotated for cross-document structure theory relationships. We will describe how we built the corpus, including our method for reducing the number of sentence pairs to be annotated by our hired judges, using lexical similarity measures. Finally, we will describe how CST and the CST Bank can be applied to different research areas such as multi-document summarization.

Keyword(s)

discourse, rhetorical structure, sentence similarity

Language(s)

English

Full Paper

411.pdf