AMQ-Bench: Agent-Oriented MCP Quality Benchmark
AMQ-Bench is an agent-oriented benchmark for evaluating the quality of Model Context Protocol (MCP) services, centered on the Repository-to-MCP-Service task. It operationalizes Availability, Usability, and Utility over a corpus of 269 repository-to-service inst
AMQ-Bench is an agent-oriented benchmark for evaluating the quality of
Model Context Protocol (MCP) services, centered on the
Repository-to-MCP-Service task. It operationalizes Availability,
Usability, and Utility over a corpus of 269 repository-to-service
instances drawn from 126 GitHub repositories across four sources
(Self-built, ToolArena, ScienceAgentBench, ResearchCodeBench). The
release includes the full corpus (amq_bench.jsonl, 269 instances), a
mini30 development subset (30 instances), a Croissant 1.0 metadata
file, and a dataset card describing the evaluation framework and
responsible-AI considerations.
This is an anonymous release accompanying a NeurIPS 2026 Evaluations
and Datasets Track submission; author and institutional information
will be added after the review period.
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.