Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

PreprintToPaper dataset

PreprintToPaper dataset: connecting bioRxiv preprints with journal publications Some preprints get published, while others remain preprints without ever being formally published in a journal. The PreprintToPaper dataset is the first of its kind, as it

👤
CreatorBadalova, Fidan
📅
Published2025-12-19
🔗
DOI10.5281/zenodo.17992421
📊
Downloads880
⚖️
Licensecc-by-4.0
File Size587.9 MB
Data TypeDataset
Published2025
Licensecc-by-4.0
Total Views2,032
Total Downloads880

PreprintToPaper dataset: connecting bioRxiv preprints with journal publications

Some preprints get published, while others remain preprints without ever being formally published in a journal. The PreprintToPaper dataset is the first of its kind, as it attempts to automatically collect publication information from bioRxiv preprints and track whether submitted preprints have resulted in a publication.

The dataset generated in this study was retrieved from the bioRxiv preprint server and the Crossref metadata API in July 2024. The dataset covers two time periods: the pre-pandemic period (2016–2018) and the COVID-19 pandemic period (2020–2022). The dataset contains detailed metadata about preprints and their published versions, including titles, authors, abstracts, institutions, submission and publication dates, licenses, and subject categories. The metadata was processed to facilitate analysis, for example, by standardizing date formats, normalizing author names, and selecting the first and last version of each preprint.

The PreprintToPaper dataset offers diverse opportunities to study changes in titles, abstracts, and author composition over different stages of a preprint, starting from the initial submission through revision stages to the final published version.

Main file provided:

  • PreprintToPaper.csv — This is the main dataset containing all processed preprints from both time periods.

Keywords

BioRxiv, Crossref, preprints, publications

Paper

A paper describing this dataset is currently under preparation.

Usage

Data Collection

As a result, 48,300 preprint records were obtained for the 2016–2018 period and 152,869 for the 2020–2022 period.

Data Processing

  • If both the initial and last versions of the preprint were available during the relevant period, both versions were retained. If only one version was available and it was the initial version, that version was retained. If only non-initial versions were available during the relevant period, the records were deleted.
  • Unlinked publications in bioRxiv were identified by checking preprints. The preprints were identified by checking preprints against two criteria: title similarity and author similarity. Tit

    📤 Share this page

    Found this useful? Share it with your network.

    ✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
PreprintToPaper dataset (Full Dataset)587.9 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Badalova, Fidan (2025). PreprintToPaper dataset. https://doi.org/10.5281/zenodo.17992421