LLM Benchmark Performance and Sentence Embeddings Dataset
Overview This dataset accompanies research on the structural similarity of large language model (LLM) performance rankings across benchmarks, as measur
Overview
This dataset accompanies research on the structural similarity of large language model (LLM) performance rankings across benchmarks, as measured through sentence-level embeddings of benchmark prompts. It contains raw performance scores for 66 frontier and open-source LLMs evaluated on four benchmarks, together with precomputed sentence embeddings produced by six embedding models under multiple chunking strategies.
All benchmark scores were collected from the HELM Capabilities leaderboard as of May 2025.
Benchmarks
Four single-turn benchmarks were selected to ensure all instances can be represented as independent text embeddings:
| Benchmark | Task Type | Score Distribution |
|---|---|---|
| MMLU-Pro | Knowledge-intensive QA | Binary per-item accuracy; smooth, high rank separability |
| GPQA | Knowledge-intensive QA | Binary per-item accuracy; smooth, high rank separability |
| IFEval | Instruction following | Discretized; reduced granularity |
| Omni-MATH | Mathematical reasoning | Bounded/discretized; reduced granularity |
</d
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.