Where Is the Automation Discourse Heading? Exploring Evolution and Trends of RPA Research Through Topic Modeling
This repository provides a comprehensive resource for bibliometric and topic modeling analyses of Robotic Process Automation (RPA) research. It includes the raw and processed data, as well as the analytical scripts used to extract and validate the 100 latent research topics. 1. Bibliometr
This repository provides a comprehensive resource for bibliometric and topic modeling analyses of Robotic Process Automation (RPA) research. It includes the raw and processed data, as well as the analytical scripts used to extract and validate the 100 latent research topics.
1. Bibliometric and Topic Modeling Data
The dataset consists of primary Excel/CSV files alongside text files for preprocessing.
data.xlsx: Includes multiple sheets containing Scopus bibliographic data, topic distributions, and derived metrics, enabling detailed exploration of thematic structures and trends.
Data from Scopus: Contains bibliographic metadata exported from Scopus, designed for use in Latent Dirichlet Allocation (LDA) analyses. Key fields include authors, abstracts, publication years, and other relevant publication information. Each paper is annotated with manually or algorithmically assigned categories and topics.
Figures 8A and 9A: Provides a detailed table of all identified topics, including associated terms, total paper counts, and citation counts per topic.
Calculation of Combined Score: Topics are ranked according to a calculated combined score, which reflects their relative importance or influence within the dataset (aggregating research interest and impact).
RPA_LDA_RESULTS_(k=100).xlsx: Contains the results of the LDA analysis with 100 topics, organized into two sheets:
top_100_terms_in_each_topic: Lists the top 100 most representative terms for each topic, facilitating interpretation of thematic content.
probabilities_with_each_topic: Provides the topic probabilities for all 3,275 documents across the 100 topics, enabling exploration of topic distributions at the document level.
Human Validation & Stratification Data (New):
topic_stratified_20%_origin.xlsx (and its CSV equivalent
filtered_abstracts_lda_sample20.csv): Contains a 20% stratified sample of the original abstracts. This dataset ensures proportional representation across all LDA topics to create a balanced dataset for human and LLM validation.finalTableLDA+HumanLabels.csv: A ground-truth validation dataset containing the sampled abstracts mapped against their algorithmically assigned LDA topics and manually annotated human labels.
Preprocessing Files: Elimination terms (extended stopwords) for the LDA analysis are provided in
stopwords.txt, and general words are listed ingeneral_words.txt.
2. Analytical Scripts and Validation Datasets
The analytical code is separated into dedicated repositories and scripts, each containing detailed e
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.