A deeply annotated testbed for geographical text analysis - Research Portal

Associated organisational units

Text available via DOI:

https://doi.org/10.1145/3149858.3149865
Final published version
Available under license: CC BY: Creative Commons Attribution 4.0 International License

A deeply annotated testbed for geographical text analysis: The Corpus of Lake District Writing

Research output: Contribution in Book/Report/Proceedings - With ISBN/ISSN › Conference contribution/Paper › peer-review

Published

Publication date	7/11/2017
Host publication	GeoHumanities'17 Proceedings of the 1st ACM SIGSPATIAL Workshop on Geospatial Humanities
Place of Publication	New York
Publisher	Association for Computing Machinery (ACM)
Pages	9-15
Number of pages	7
ISBN (print)	9781450354967
<mark>Original language</mark>	English

Abstract

This paper describes the development of an annotated corpus which forms a challenging testbed for geographical text analysis methods. This dataset, the Corpus of Lake District Writing (CLDW), consists of 80 manually digitised and annotated texts (comprising over 1.5 million word tokens). These texts were originally composed between 1622 and 1900, and they represent a range of different genres and authors. Collectively, the texts in the CLDW constitute an indicative sample of writing about the English Lake District during the early seventeenth century and the early twentieth century. The corpus is annotated more deeply than is currently possible with vanilla Named Entity Recognition, Disambiguation and geoparsing. This is especially true of the geographical information the corpus contains, since we have undertaken not only to link different historical and spelling variants of place-names, but also to identify and to differentiate geographical features such as waterfalls, woodlands, farms or inns. In addition, we illustrate the potential of the corpus as a gold standard by evaluating the results of three different NLP libraries and geoparsers on its contents. In the evaluation, the standard NER processing of the text by the different NLP libraries produces many false positive and false negative results, showing the strength of the gold standard.

Research

Associated organisational units

Links

Text available via DOI:

A deeply annotated testbed for geographical text analysis: The Corpus of Lake District Writing

Abstract

Quick Links

Connect With Us

Faculties & Depts

Contact Us

Research

Associated organisational units

Links

Text available via DOI:

A deeply annotated testbed for geographical text analysis: The Corpus of Lake District Writing

Abstract

Related research outputs

Geospatial Innovation in the Digital Humanities: Implementation and Evaluation of Deep Mapping in the Lake District

Quick Links

Connect With Us

Faculties & Depts

Contact Us