MockEval: A Benchmark for Evaluating LLMs on Mock-Aware Unit Testing
# IntroductionThis is the replication package for the paper "MockEval: A Benchmark for Evaluating LLMs on Mock-AwareUnit Testing". ## Task DefinitionEvaluate the ability of large language models in identifying, classifying, and generating mock code, given the focal method and c
# Introduction
This is the replication package for the paper “MockEval: A Benchmark for Evaluating LLMs on Mock-Aware
Unit Testing”.
## Task Definition
Evaluate the ability of large language models in identifying, classifying, and generating mock code, given the focal method and context.
# Setup
“`bash
pip install -r requirement.txt
cd construction
“`
## The static of dataset
MockEval benchmark:
Database/ORM 32
File/IO 93
Internal/Logic 577
ML/DATA 70
Network/HTTP 201
System/Process 74
Mock Cases 900
Mock-Free Cases 804
# Pipline
## Benchmark Construction
1. Repository Selection
node1 server
OUTPUT:candidates.json
“`bash
python 1_search_repos.py
python 2_clone_and_scan.py
“`
2. Mock Case Extract
INPUT:candidates.json
OUTPUT:valid_mock_projects.json
“`bash
python 3_fast_retry.py
python 4_clean_small.py
python 5_extract_pairs.py
“`
3. Focal Method Matching
INPUT:benchmark_data_pairs.json
OUTPUT:benchmark_dataset_final.json
“`bash
python 6_locate_focal_method.py
“`
4. Context Extract
INPUT:benchmark_dataset_final.json
OUTPUT:mock_benchmark_context_enriched.json,merged_benchmark.json
“`bash
python 8_enrich_context_v2.py
python 9_compare_context.py
“`
5. Mock-Free Case Selection
INPUT:mock_benchmark_high_quality.json
OUTPUT:negative_samples_benchmark.json
“`bash
python collect_negative_samples.py
python generate_negative_samples.py
“`
6. Human Annotation
INPUT:mock_benchmark_high_quality.json,negative_samples_benchmark.json
OUTPUT:cleaned_benchmark.json
“`bash
python collect_final_dataset_v3.py
python merge.py
“`
## Experiment
1. Mock Necessity & Type
INPUT: cleaned_benchmark.json
OUTPUT: results_coder_v2_lite_classification.json,results_r1_70b_classification.json,results_r1_32b_classification.json,results_r1_14b_classification.json,gemini3_pro_classification.json,gpt5_2_results_classification.json,results_14b_classification.json,results_32b_classification.json,results_30b_a3b_classification.json,claude_opus_4_5_multistage_progress_classification.json
“`bash
python run_14b_classification.py
python run_32b_classification.py
python run_30b_a3b_classification.py
python run_claude_opus_4_5.py
python run_deepseek_r1_14b.py
python run_deepseek_r1_32b.py
python run_deepseek_r1_70b.py
python run_gemini3_pro.py
python run_gpt5_2.py
python run_coder_v2_lite.py
“`
2. Mock Generation
INPUT: cleaned_benchmark.json
OUTPUT: results_coder_v2_lite_gen_stages.json,results_r1_70b_gen_stages.json,results_r1_32b_gen_stages.json,results_r1_14b_gen_stages.json,results_gemini_mock_gen.json,gpt5_2_mock_gen_results.json,results_qwen14b_gen_stages.json,results_qwen32b_gen_stages.json,results_30b_gen_stages.json
“`bash
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.