Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Webis-TLDR-17 Corpus

This corpus contains preprocessed posts from the Reddit dataset, suitable for abstractive summarization using deep learning. The format is a json file where each line is a JSON object representing a post. The schema of each post is shown below: author: string (nullable = true)

👤
CreatorSyed, Shahbaz
📅
Published2017-11-07
🔗
DOI10.5281/zenodo.1043504
📊
Downloads22,378
⚖️
Licensecc-by-4.0
File Size2.9 GB
Data TypeDataset
Published2017
Licensecc-by-4.0
Total Views7,487
Total Downloads22,378

This corpus contains preprocessed posts from the Reddit dataset, suitable for abstractive summarization using deep learning. The format is a json file where each line is a JSON object representing a post. The schema of each post is shown below:

  • author: string (nullable = true)
  • body: string (nullable = true)
  • normalizedBody: string (nullable = true)
  • content: string (nullable = true)
  • content_len: long (nullable = true)
  • summary: string (nullable = true)
  • summary_len: long (nullable = true)
  • id: string (nullable = true)
  • subreddit: string (nullable = true)
  • subreddit_id: string (nullable = true)
  • title: string (nullable = true)

Specifically, the content and summary fields can be directly used as inputs to a deep learning model (e.g. Sequence to Sequence model ). The dataset consists of 3,848,330 posts with an average length of 270 words for content, and 28 words for the summary. The dataset is a combination of both the Submissions and Comments merged on the common schema. As a result, most of the comments which do not belong to any submission have null as their title.

Note : This corpus does not contain a separate test set. Thus it is up to the users to divide the corpus into appropriate training, validation and test sets.

 

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Webis-TLDR-17 Corpus (Full Dataset)2.9 GB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Syed, Shahbaz (2017). Webis-TLDR-17 Corpus. https://doi.org/10.5281/zenodo.1043504