Lexical Substitution Dataset for German.

Cholakov, Kostadin; Biemann, Chris; Eckle-Kohler, Judith; Gurevych, Iryna

Open Access

Lexical Substitution Dataset for German.

Files

germeval2015.tar.xz (360.87 KB)

Date

2020-07-24

Type

Dataset

Authors

Cholakov, Kostadin
Biemann, Chris

Eckle-Kohler, Judith
Gurevych, Iryna

Description

This article describes a lexical substitution dataset for German. The whole dataset contains 2,040 sentences from the German Wikipedia,with one target word in each sentence. There are 51 target nouns, 51 adjectives, and 51 verbs randomly selected from 3 frequency groups based on the lemma frequency list of the German WaCKy corpus. 200 sentences have been annotated by 4 professional annotators and the remaining sentences by 1 professional annotator and 5 additional annotators who have been recruited via crowdsourcing. The resulting dataset can be used to evaluate not only lexical substitution systems, but also different sense inventories and word sense disambiguation systems.

Identifier

https://tudatalib.ulb.tu-darmstadt.de/handle/tudatalib/2436

Related Resources

Is Supplement To

http://www.lrec-conf.org/proceedings/lrec2014/pdf/545_Paper

DFG Classification

4.43-04 Künstliche Intelligenz und Maschinelle Lernverfahren
4.43-05 Bild- und Sprachverarbeitung, Computergraphik und Visualisierung, Human Computer Interaction, Ubiquitous und Wearable Computing

Collections

Lexical Substitution

License

Except where otherwise noted, this license is described as CC BY-SA 3.0

Full item page

Lexical Substitution Dataset for German.

Files

Date

Type

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Description

Keywords

Citation

Identifier

Endorsement

Related Resources

Is Supplement To

DFG Classification

Project(s)

Faculty

Collections

License