Replication Package: Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering
This replication package contains the data and scripts used to generate and evaluate LLM-generated threats to validity for 375 ICSE research papers. Contents: Input Papers (Input_Papers.zip) - The dataset of 375 ICSE papers used a
This replication package contains the data and scripts used to generate and evaluate LLM-generated threats to validity for 375 ICSE research papers.
Contents:
Input Papers (Input_Papers.zip) – The dataset of 375 ICSE papers used as input for threat generation.
Threat Generation Prompt (threat_generation_prompt.txt) – The prompt used to generate threats to validity for all 375 papers. Each paper was passed (excluding its existing threats to validity section) to generate_threats.ipynb alongside this prompt.
Generated Threats Dataset (threat_dataset.csv) – The LLM-generated threats for all 375 papers. Note: The dataset contains 110 entries rather than 375 because the generation script occasionally combines multiple papers into a single row when Gemini produces a low number of threats for individual papers.
Scripts:
paper_split.py: Splits papers (separating the TTV section) for processing.
generate_threats.ipynb: Notebook for running the threat generation across all papers.
generate_rubric.py: LLM-as-judge script that grades generated threats against a 4-criterion rubric (Relevance, Specificity, Clarity, Mitigation).
rubric_count.py: Counts how many threats fall into each scoring bucket per criterion (e.g., how many threats had High Specificity, how many had No Impact for Relevance).
compare_threats.py: Semantically compares LLM-generated threats against the paper’s actual TTV section using Gemini 2.5 Flash. Buckets each threat into matched, only_AI, or only_Paper across the three validity categories (external, internal, construct). matched threats appear in both the prediction and the paper, only_AI threats were predicted but not in the paper, and only_Paper threats were in the paper but missed by the AI.
Generated Rubric (evaluation_rubric.txt) – The rubric used by generate_rubric.py to grade each predicted threat across the four criteria (Relevance, Specificity, Clarity, Mitigation).
Scoring Sample (scoring.csv) – A sample of manually graded threats used to establish inter-rater agreement for the rubric criteria.
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.