HF-NLP10K: A Dataset and Metadata Analysis of 10,000+ Hugging Face NLP Models
Hugging Face has become a central platform for sharing NLP models. However, inconsistent and incomplete metadata makes it difficult to find suitable models. In this project, we introduce HF-NLP10K, a structured dataset of over 10,000 NLP models, including many
Hugging Face has become a central platform for sharing NLP models. However, inconsistent and incomplete metadata makes it difficult to find suitable models. In this project, we introduce HF-NLP10K, a structured dataset of over 10,000 NLP models, including many LLMs, enriched with 33 metadata fields, combining Hugging Face API data with LLM-based parsing of free-text model cards.
Here we provide:
- 📦 The HF-NLP10K dataset (
CMA-PROJ_Dataset.csv.gz) - 📜 Scripts for dataset creation
Primary contact:
Michael Färber (TU Dresden and ScaDS.AI, Germany) – michael.faerber@tu-dresden.de
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.