Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Replication Package: Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering

This replication package contains the data and scripts used to generate and evaluate LLM-generated threats to validity for 375 ICSE research papers. Contents: Input Papers (Input_Papers.zip) - The dataset of 375 ICSE papers used a

👤
CreatorDang, Ryan
📅
Published2026-05-29
🔗
DOI10.5281/zenodo.20453824
📊
Downloads12
⚖️
Licensecc-by-4.0
File Size364.1 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views78
Total Downloads12

This replication package contains the data and scripts used to generate and evaluate LLM-generated threats to validity for 375 ICSE research papers.

Contents:

  1. Input Papers (Input_Papers.zip) – The dataset of 375 ICSE papers used as input for threat generation.

  2. Threat Generation Prompt (threat_generation_prompt.txt) – The prompt used to generate threats to validity for all 375 papers. Each paper was passed (excluding its existing threats to validity section) to generate_threats.ipynb alongside this prompt.

  3. Generated Threats Dataset (threat_dataset.csv) – The LLM-generated threats for all 375 papers. Note: The dataset contains 110 entries rather than 375 because the generation script occasionally combines multiple papers into a single row when Gemini produces a low number of threats for individual papers.

  4. Scripts:

    • paper_split.py: Splits papers (separating the TTV section) for processing.

    • generate_threats.ipynb: Notebook for running the threat generation across all papers.

    • generate_rubric.py: LLM-as-judge script that grades generated threats against a 4-criterion rubric (Relevance, Specificity, Clarity, Mitigation).

    • rubric_count.py: Counts how many threats fall into each scoring bucket per criterion (e.g., how many threats had High Specificity, how many had No Impact for Relevance).

    • compare_threats.py: Semantically compares LLM-generated threats against the paper’s actual TTV section using Gemini 2.5 Flash. Buckets each threat into matched, only_AI, or only_Paper across the three validity categories (external, internal, construct). matched threats appear in both the prediction and the paper, only_AI threats were predicted but not in the paper, and only_Paper threats were in the paper but missed by the AI.

  5. Generated Rubric (evaluation_rubric.txt) – The rubric used by generate_rubric.py to grade each predicted threat across the four criteria (Relevance, Specificity, Clarity, Mitigation). 

  6. Scoring Sample (scoring.csv) – A sample of manually graded threats used to establish inter-rater agreement for the rubric criteria.

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Replication Package: Do LLMs Understand Validity? An Empirical… (Full Dataset)364.1 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Dang, Ryan (2026). Replication Package: Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering. https://doi.org/10.5281/zenodo.20453824