Curtotti and McCreath, ‘A Corpus of Australian Contract Language: Description, Profiling and Analysis’

Abstract:
Written contracts are a fundamental framework for economic and cooperative transactions in society. Little work has been reported on the application of natural language processing or corpus linguistics to contracts. In this paper we report the design, pro ling and initial analysis of a corpus of Australian contract language. This corpus enables a quantitative and qualitative characterisation of Australian contract language as an input to the development of contract drafting tools. Profi ling of the corpus is consistent with its suitability for use in language engineering applications. We provide descriptive statistics for the corpus and show that document length and document vocabulary size approximate to log normal distributions. The corpus conforms to Zipf’s law and comparative type to token ratios are consistent with lower term sparsity (an expectation for legal language). We highlight distinctive term usage in Australian contract language. Results derived from the corpus indicate a longer prepositional phrase depth in sentences in contract rules extracted from the corpus, as compared to other corpora.

Curtotti, Michael and McCreath, Eric, A Corpus of Australian Contract Language: Description, Profiling and Analysis (June 6, 2011). Proceedings of the 13th International Conference on Artificial Intelligence and Law. ACM, 2011.

First posted 2013-08-02 13:41:50

Leave a Reply