Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Understanding and Improving LLM-Based Test Oracle Generation

Replication Package This is the replication package for ASE submission, containing both scripts and data that are requested by the replication. It also provides detailed instructions to replicate the evaluation. This package contains the data, prompts, scripts, and model outputs u

👤
CreatorAnonymous
📅
Published2026-03-26
🔗
DOI10.5281/zenodo.19228870
📊
Downloads7
⚖️
Licensecc-by-4.0
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views117
Total Downloads7

Replication Package

This is the replication package for ASE submission, containing both scripts and data that are requested by the replication. It also provides detailed instructions to replicate the evaluation.

This package contains the data, prompts, scripts, and model outputs used for assertion-generation evaluation on Defects4J and GHRB test cases.

1. Directory Layout

  • data/

    • dataset.jsonl: main testcase-level dataset used by GPT-4o and MiniMax scripts.

    • toga/

      • data_input.csv

      • data_meta.csv

    • togll/

      • dataset.jsonl

  • prompts/

    • old_prompt.txt: original prompt.

    • new_prompt.txt: knowledge-augmented prompt.

  • scripts/

    • eval_assertion_with_gpt4o_old.py

    • eval_assertion_with_gpt4o_new.py

    • eval_assertion_with_minimax_old_new.py

    • run_toga.py

    • run_togll.py

  • outputs/

    • gpt-4o/old/: one text file per testcase id.

    • gpt-4o/new/: one text file per testcase id.

    • minimax/old/: one text file per testcase id.

    • minimax/new/: one text file per testcase id.

    • toga/result.csv: TOGA oracle output.

    • togll/result.csv: TOGLL oracle output.

    • togll/result.json: TOGLL inference raw result.

    • result matrix.csv: consolidated status matrix.

2. Environment

  • Python 3.9+

  • Install dependencies (example):

pip install requests pandas torch transformers

Note: TOGLL/TOGA scripts may require additional model files/checkpoints that are not bundled in this folder.

3. Running GPT-4o Evaluation

Before running, edit API fields in scripts:

  • API_KEY = "YOUR API KEY HERE"

  • API_URL = "https://api.openai.com/v1"

Run:

python scripts/eval_assertion_with_gpt4o_old.py
python scripts/eval_assertion_with_gpt4o_new.py

Default runtime outputs are written under:

  • outputs/gpt-4o/old_runtime/

  • outputs/gpt-4o/new_runtime/

4. Running MiniMax Evaluation

Edit in scripts/eval_assertion_with_minimax_old_new.py:

📦
Understanding and Improving LLM-Based Test Oracle Generation (Full Dataset)Size varies
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Anonymous (2026). Understanding and Improving LLM-Based Test Oracle Generation. https://doi.org/10.5281/zenodo.19228870