Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

LLM Benchmark Performance and Sentence Embeddings Dataset

Overview This dataset accompanies research on the structural similarity of large language model (LLM) performance rankings across benchmarks, as measur

👤
CreatorAnonymous
📅
Published2026-05-25
🔗
DOI10.5281/zenodo.20384643
📊
Downloads11
⚖️
Licensecc-by-4.0
File Size642.0 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views66
Total Downloads11

Overview

This dataset accompanies research on the structural similarity of large language model (LLM) performance rankings across benchmarks, as measured through sentence-level embeddings of benchmark prompts. It contains raw performance scores for 66 frontier and open-source LLMs evaluated on four benchmarks, together with precomputed sentence embeddings produced by six embedding models under multiple chunking strategies.

All benchmark scores were collected from the HELM Capabilities leaderboard as of May 2025.

Benchmarks

Four single-turn benchmarks were selected to ensure all instances can be represented as independent text embeddings:

BenchmarkTask TypeScore Distribution
MMLU-ProKnowledge-intensive QABinary per-item accuracy; smooth, high rank separability
GPQAKnowledge-intensive QABinary per-item accuracy; smooth, high rank separability
IFEvalInstruction followingDiscretized; reduced granularity
Omni-MATHMathematical reasoningBounded/discretized; reduced granularity

</d

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
LLM Benchmark Performance and Sentence Embeddings Dataset (Full Dataset)642.0 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Anonymous (2026). LLM Benchmark Performance and Sentence Embeddings Dataset. https://doi.org/10.5281/zenodo.20384643