Open Access

Wikipedia Text Segmentation

Files

WikipediaTextSegmentation.zip (10.98 MB)

Date

2020-07-25

Type

Dataset

Authors

Martin, Marko
Zesch, Torsten
Erbs, Nicolai
Gurevych, Iryna

Description

For corpus generation, we extracted top-level sections of featured articles and concatenated their textual contents to a pure-text corpus file. The content of a section is constituted by the concatenation of the text of its paragraph elements and the content of contained sections. Particularly, other elements such as tables and image captions are ignored during generating the text for a section because text segmentation is meant to be applied to prose and not to pieces of information such as table fields. Furthermore, sections with one of the titles ``See also'', ``References'', and ``External links'' are skipped as they do not contain information where segmentation makes sense.

Identifier

https://tudatalib.ulb.tu-darmstadt.de/handle/tudatalib/2454

DFG Classification

4.43-04 Künstliche Intelligenz und Maschinelle Lernverfahren
4.43-05 Bild- und Sprachverarbeitung, Computergraphik und Visualisierung, Human Computer Interaction, Ubiquitous und Wearable Computing

Collections

Text Segmentation

License

Except where otherwise noted, this license is described as CC BY-SA 3.0

Full item page

Wikipedia Text Segmentation

Files

Date

Type

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Description

Keywords

Citation

Identifier

Endorsement

DFG Classification

Project(s)

Faculty

Collections

License