Zur Kurzanzeige

dc.contributor.authorCholakov, Kostadin
dc.contributor.authorBiemann, Chris
dc.contributor.authorEckle-Kohler, Judith
dc.contributor.authorGurevych, Iryna
dc.date.accessioned2020-07-24T12:49:24Z
dc.date.available2020-07-24T12:49:24Z
dc.date.issued2014
dc.identifier.urihttps://tudatalib.ulb.tu-darmstadt.de/handle/tudatalib/2436
dc.descriptionThis article describes a lexical substitution dataset for German. The whole dataset contains 2,040 sentences from the German Wikipedia,with one target word in each sentence. There are 51 target nouns, 51 adjectives, and 51 verbs randomly selected from 3 frequency groups based on the lemma frequency list of the German WaCKy corpus. 200 sentences have been annotated by 4 professional annotators and the remaining sentences by 1 professional annotator and 5 additional annotators who have been recruited via crowdsourcing. The resulting dataset can be used to evaluate not only lexical substitution systems, but also different sense inventories and word sense disambiguation systems.en_US
dc.relationIsSupplementTo;URL;http://www.lrec-conf.org/proceedings/lrec2014/pdf/545_Paper
dc.rightsCC BY-SA 3.0
dc.rights.urihttps://creativecommons.org/licenses/by-sa/3.0/
dc.subject.classification4.43-04 Künstliche Intelligenz und Maschinelle Lernverfahrenen_US
dc.subject.classification4.43-05 Bild- und Sprachverarbeitung, Computergraphik und Visualisierung, Human Computer Interaction, Ubiquitous und Wearable Computing
dc.subject.ddc004
dc.titleLexical Substitution Dataset for German.en_US
dc.typeDataseten_US
tud.history.classificationVersion=2016-2020;409-05 Interaktive und intelligente Systeme, Bild- und Sprachverarbeitung, Computergraphik und Visualisierung


Dateien zu dieser Ressource

Thumbnail

Der Datensatz erscheint in:

Zur Kurzanzeige

CC BY-SA 3.0
Solange nicht anders angezeigt, wird die Lizenz wie folgt beschrieben: CC BY-SA 3.0