TUdatalib : Lexical Substitution Dataset for German.

Zur Kurzanzeige

dc.contributor.author	Cholakov, Kostadin
dc.contributor.author	Biemann, Chris
dc.contributor.author	Eckle-Kohler, Judith
dc.contributor.author	Gurevych, Iryna
dc.date.accessioned	2020-07-24T12:49:24Z
dc.date.available	2020-07-24T12:49:24Z
dc.date.issued	2014
dc.identifier.uri	https://tudatalib.ulb.tu-darmstadt.de/handle/tudatalib/2436
dc.description	This article describes a lexical substitution dataset for German. The whole dataset contains 2,040 sentences from the German Wikipedia,with one target word in each sentence. There are 51 target nouns, 51 adjectives, and 51 verbs randomly selected from 3 frequency groups based on the lemma frequency list of the German WaCKy corpus. 200 sentences have been annotated by 4 professional annotators and the remaining sentences by 1 professional annotator and 5 additional annotators who have been recruited via crowdsourcing. The resulting dataset can be used to evaluate not only lexical substitution systems, but also different sense inventories and word sense disambiguation systems.	en_US
dc.relation	IsSupplementTo;URL;http://www.lrec-conf.org/proceedings/lrec2014/pdf/545_Paper
dc.rights	CC BY-SA 3.0
dc.rights.uri	https://creativecommons.org/licenses/by-sa/3.0/
dc.subject.classification	4.43-04 Künstliche Intelligenz und Maschinelle Lernverfahren	en_US
dc.subject.classification	4.43-05 Bild- und Sprachverarbeitung, Computergraphik und Visualisierung, Human Computer Interaction, Ubiquitous und Wearable Computing
dc.subject.ddc	004
dc.title	Lexical Substitution Dataset for German.	en_US
dc.type	Dataset	en_US
tud.history.classification	Version=2016-2020;409-05 Interaktive und intelligente Systeme, Bild- und Sprachverarbeitung, Computergraphik und Visualisierung

Dateien zu dieser Ressource

Name:: germeval2015.tar.xz
Größe:: 360.8KB
Format:: Unbekannt

Anzahl der Dateien

Der Datensatz erscheint in:

Lexical Substitution [3]

Zur Kurzanzeige

Solange nicht anders angezeigt, wird die Lizenz wie folgt beschrieben: CC BY-SA 3.0