Online Appendix – Making Sense of AI Agents Hype: Adoption, Architectures, and Takeaways from Practitioners
This replication package provides how we collect, extract, code, and analyze findings from Videos (transcripts) related to the AI agent-based. In this research, the LLMs process was operated on Finnish Supercomputer using vLLM of LRM (Large Reasoning Model
This replication package provides how we collect, extract, code, and analyze findings from Videos (transcripts) related to the AI agent-based.
In this research, the LLMs process was operated on Finnish Supercomputer using vLLM of LRM (Large Reasoning Models) and LLMs (Large Language Models).
Note: All the hyperlinks referring to the shared files only work in the local version, as the browser does not find the files. (Please download the replication package here.)
License
All generated data is provided under DATA_LICENSE Creative Commons 4.0 Attribution License.
All scripts are provided under the Script_LICENSE MIT License.
Contents
This repository consists of the following files:
README.md: The README file describes the project and how to run the scripts to get the data and do the analysis.
INSTALL.md: Installation and configurations instructions.
requirements.txt: Library requirements for the Python environment.
Data_Collection
- Analysis.xlsx: File to describe the process of the data selection. Among tabs, Tab ’Talks Selection’ describes selection of talks by application of the I/E criteria by 4 human assessors. Tab ‘I/E Criteria’ is the Inclusion and Exclusion Criteria used for the Talk Selection.
Scripts
00_Data_Collection
- audioCutting.py: Script to cutting the audio.
- audioExtraction.sh: Script to extract the audio.
- transcriptExtraction.sh: Script to extract the transcripts.
- videoDownload.sh: Script to download the videos.
- Transcripts_Pipeline_Description: General description of the transcript extraction pipeline.
01_Extraction
01_videolinkscript_extraction.py: Script to extract evidence-grounded answers to 7 research questions from transcripts using retrieval-augmented generation and SWEBOK-aligned terminology refinement.
01_sbatch_extraction.sh: Script to submit the extraction work to the supercomputer.
02_Extraction_Validate
02_extraction_validate.py: Script to validate LLM-generated answers to research questions against transcript evidence using retrieval-augmented context and structured judgment prompts.
02_extraction_validate.sh: Script to submit the extraction validation work to the supercomputer.
03_ThematicAxial
03_thematicaxial_RQ1.py: Script to perform thematic and axial coding of RQ1-related excerpts by synthesizing timelines, motivations, an
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.