Writing the Vernacular: Transcribing and Tagging the Newcastle Electronic Corpus of Tyneside English

Home > Research > Publications & Outputs > Writing the Vernacular: Transcribing and Taggin...

Computing and Communications

Associated organisational units

Keywords

cs_eprint_id, 1518 cs_uid, 355

View graph of relations

Research output: Contribution to Journal/Magazine › Journal article › peer-review

Published

J. Beal
K. Corrigan
N. Smith
P. Rayson

More...

<mark>Journal publication date</mark>	2007
<mark>Journal</mark>	Studies in Variation, Contacts and Change in English
Volume	1
Publication Status	Published
<mark>Original language</mark>	English

Abstract

The Newcastle Electronic Corpus of Tyneside English (NECTE) presented a number of problems not encountered by those producing corpora of standard varieties. The primary material consisted of audio recordings which needed to be orthographically transcribed and grammatically tagged. Preston (1985), (2000), Macaulay (1991), Kirk (1997), Cameron (2001) and Beal (2005) all note that representing vernacular Englishes orthographically, e.g. by using "eye dialect", can be problematic on various levels. Apart from unwelcome associations with negative political, racial or social connotations, there are theoretical objections to devising non-standard spellings which represent certain groups of vernacular speakers, thus making their speech appear more differentiated from mainstream colloquial varieties than is warranted. In the first half of this paper, we outline the principles and methods adopted in devising an Orthographic Transcription Protocol (OTP) for such a vernacular corpus, and the challenges faced by the NECTE team in practice. Protocols for grammatical tagging have likewise been devised with standard varieties in mind. In the second half, we relate how existing part-of-speech (POS)-tagging software (CLAWS4, cf. Garside & Smith 1997; and Template Tagger, cf. Fligelstone et al. 1997) had to be adapted to take account of the non-standard grammar of Tyneside English.

Research

Associated organisational units

Links

Keywords

Writing the Vernacular: Transcribing and Tagging the Newcastle Electronic Corpus of Tyneside English

Abstract

Quick Links

Connect With Us

Faculties & Depts

Contact Us