Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

HF-NLP10K: A Dataset and Metadata Analysis of 10,000+ Hugging Face NLP Models

Hugging Face has become a central platform for sharing NLP models. However, inconsistent and incomplete metadata makes it difficult to find suitable models. In this project, we introduce HF-NLP10K, a structured dataset of over 10,000 NLP models, including many

👤
CreatorFärber, Michael
📅
Published2025-06-17
🔗
DOI10.5281/zenodo.15682522
📊
Downloads2,628
⚖️
Licensecc-by-4.0
File Size50.1 MB
Data TypeDataset
Published2025
Licensecc-by-4.0
Total Views2,799
Total Downloads2,628

Hugging Face has become a central platform for sharing NLP models. However, inconsistent and incomplete metadata makes it difficult to find suitable models. In this project, we introduce HF-NLP10K, a structured dataset of over 10,000 NLP models, including many LLMs, enriched with 33 metadata fields, combining Hugging Face API data with LLM-based parsing of free-text model cards.

Here we provide:

  • 📦 The HF-NLP10K dataset (CMA-PROJ_Dataset.csv.gz)
  • 📜 Scripts for dataset creation

Primary contact:

Michael Färber (TU Dresden and ScaDS.AI, Germany) – michael.faerber@tu-dresden.de

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
HF-NLP10K: A Dataset and Metadata Analysis of 10,000+… (Full Dataset)50.1 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Färber, Michael (2025). HF-NLP10K: A Dataset and Metadata Analysis of 10,000+ Hugging Face NLP Models. https://doi.org/10.5281/zenodo.15682522