Featured Dataset
Uzbek 65 million web corpus
A contemporary web corpus gathered in 2026. The corpus contains approximately 60 million tokens. <t
File Size507.3 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views123
Total Downloads90
A contemporary web corpus gathered in 2026.
The corpus contains approximately 60 million tokens.
| Name of the corpus | Number of tokens | Number of sentences |
| Wikipedia resource | 27711575 | 2481232 |
| Web resources | 33169827 | 2389912 |
| School Corpus | 1408830 | 154239 |
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.
Khajibaeva, Surayyo (2026). Uzbek 65 million web corpus. https://doi.org/10.5281/zenodo.19462612