Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Webis-Web-Archive-17

The Webis-Web-Archive-17 comprises a total of 10,000 web page archives from mid-2017 that were carefully sampled from the Common Crawl to involve a mixture of high-ranking and low-ranking web pages. The dataset contains the web archive files, HTML DOM, and screenshots of each web page, as well as

👤
CreatorKiesel, Johannes
📅
Published2017-10-04
🔗
DOI10.5281/zenodo.4064019
📊
Downloads824,058
⚖️
Licensecc-by-sa-4.0
File Size92.7 GB
Data TypeDataset
Published2017
Licensecc-by-sa-4.0
Total Views3,298
Total Downloads824,058

The Webis-Web-Archive-17 comprises a total of 10,000 web page archives from mid-2017 that were carefully sampled from the Common Crawl to involve a mixture of high-ranking and low-ranking web pages. The dataset contains the web archive files, HTML DOM, and screenshots of each web page, as well as per-page annotations of visual web archive quality. See this overview for all datasets that built upon this one. If you use this dataset in your research, please cite it using this paper.

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Webis-Web-Archive-17 (Full Dataset)92.7 GB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Kiesel, Johannes (2017). Webis-Web-Archive-17. https://doi.org/10.5281/zenodo.4064019