Home > Research > Publications & Outputs > A New Aligned Simple German Corpus

Links

Text available via DOI:

View graph of relations

A New Aligned Simple German Corpus

Research output: Contribution in Book/Report/Proceedings - With ISBN/ISSNConference contribution/Paperpeer-review

Published
  • Vanessa Toborek
  • Moritz Busch
  • Malte Boßert
  • Christian Bauckhage
  • Pascal Welke
Close
Publication date9/07/2023
Host publicationProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Place of PublicationStroudsburg, PA
PublisherAssociation for Computational Linguistics (ACL Anthology)
Pages11393-11412
Number of pages20
ISBN (electronic)9781959429722
<mark>Original language</mark>English

Abstract

“Leichte Sprache”, the German counterpart to Simple English, is a regulated language aiming to facilitate complex written language that would otherwise stay inaccessible to different groups of people. We present a new sentence-aligned monolingual corpus for Simple German – German. It contains multiple document-aligned sources which we have aligned using automatic sentence-alignment methods. We evaluate our alignments based on a manually labelled subset of aligned documents. The quality of our sentence alignments, as measured by the F1-score, surpasses previous work. We publish the dataset under CC BY-SA and the accompanying code under MIT license.

Bibliographic note

DBLP's bibliographic metadata records provided through http://dblp.org/search/publ/api are distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Although the bibliographic metadata records are provided consistent with CC0 1.0 Dedication, the content described by the metadata records is not. Content may be subject to copyright, rights of privacy, rights of publicity and other restrictions.