Wikipedia Text Segmentation
Loading...
Date
2020-07-25
Type
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Description
For corpus generation, we extracted top-level sections of featured articles and concatenated their textual contents to a pure-text corpus file. The content of a section is constituted by the concatenation of the text of its paragraph elements and the content of contained sections. Particularly, other elements such as tables and image captions are ignored during generating the text for a section because text segmentation is meant to be applied to prose and not to pieces of information such as table fields. Furthermore, sections with one of the titles ``See also'', ``References'', and ``External links'' are skipped as they do not contain information where segmentation makes sense.
Keywords
Citation
Endorsement
Project(s)
Faculty
Collections
License
Except where otherwise noted, this license is described as CC BY-SA 3.0