Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Uzbek 65 million web corpus

A contemporary web corpus gathered in 2026. The corpus contains approximately 60 million tokens. <t

👤
CreatorKhajibaeva, Surayyo
📅
Published2026-04-07
🔗
DOI10.5281/zenodo.19462612
📊
Downloads90
⚖️
Licensecc-by-4.0
File Size507.3 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views123
Total Downloads90

A contemporary web corpus gathered in 2026.

The corpus contains approximately 60 million tokens.

Name of the corpusNumber of tokensNumber of sentences
Wikipedia resource277115752481232
Web resources331698272389912
School Corpus1408830154239

 

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Uzbek 65 million web corpus (Full Dataset)507.3 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Khajibaeva, Surayyo (2026). Uzbek 65 million web corpus. https://doi.org/10.5281/zenodo.19462612