Lexical Substitution Dataset for German.

datacite.relation.isSupplementTo http://www.lrec-conf.org/proceedings/lrec2014/pdf/545_Paper
dc.contributor.author Cholakov, Kostadin
dc.contributor.author Biemann, Chris
dc.contributor.author Eckle-Kohler, Judith
dc.contributor.author Gurevych, Iryna
dc.date.accessioned 2020-07-24T12:49:24Z
dc.date.available 2020-07-24T12:49:24Z
dc.date.created 2014
dc.date.issued 2020-07-24
dc.description This article describes a lexical substitution dataset for German. The whole dataset contains 2,040 sentences from the German Wikipedia,with one target word in each sentence. There are 51 target nouns, 51 adjectives, and 51 verbs randomly selected from 3 frequency groups based on the lemma frequency list of the German WaCKy corpus. 200 sentences have been annotated by 4 professional annotators and the remaining sentences by 1 professional annotator and 5 additional annotators who have been recruited via crowdsourcing. The resulting dataset can be used to evaluate not only lexical substitution systems, but also different sense inventories and word sense disambiguation systems. en_US
dc.identifier.uri https://tudatalib.ulb.tu-darmstadt.de/handle/tudatalib/2436
dc.rights CC BY-SA 3.0
dc.rights.licenseother
dc.rights.uri https://creativecommons.org/licenses/by-sa/3.0/
dc.subject.classification 4.43-04
dc.subject.classification 4.43-05
dc.subject.ddc 004
dc.title Lexical Substitution Dataset for German. en_US
dc.type Dataset en_US
dcterms.accessRights openAccess
person.identifier.orcid #PLACEHOLDER_PARENT_METADATA_VALUE#
person.identifier.orcid 0000-0002-8449-9624
person.identifier.orcid #PLACEHOLDER_PARENT_METADATA_VALUE#
person.identifier.orcid #PLACEHOLDER_PARENT_METADATA_VALUE#
tuda.history.classification Version=2016-2020;409-05 Interaktive und intelligente Systeme, Bild- und Sprachverarbeitung, Computergraphik und Visualisierung

Files

Original bundle

Now showing 1 - 1 of 1
NameDescriptionSizeFormat
germeval2015.tar.xz360.87 KBUnknown data format Download