AIClusterKPI: A Multivariate KPI Dataset for Cloud-Native Anomaly Detection
Overview This dataset provides a comprehensive collection of multivariate Key Performance Indicator (KPI) traces covering a continuous period of 16 days with a temporal resolution of one minute. It is specifically designed for time-series anomaly detection models in cloud-native environm
Overview
This dataset provides a comprehensive collection of multivariate Key Performance Indicator (KPI) traces covering a continuous period of 16 days with a temporal resolution of one minute. It is specifically designed for time-series anomaly detection models in cloud-native environments.
Dataset Composition
The 31 features are categorized into two main groups:
Node-Level Metrics (28 dimensions): For each of the 7 monitored nodes, four key performance indicators (KPIs) are captured:
CPU Utilization: Calculated as a percentage (%).
Available Memory: Measured in Gigabytes (GB).
Disk I/O Wait Time: Representing I/O pressure in seconds per second (s/s).
Network Inbound Traffic: Measured in Megabytes per second (MB/s).
Cluster-Level Metrics (3 dimensions): These provide a global view of the cluster health:
Pending Pods: The total count of pods awaiting scheduling.
API Server Latency: The 99th percentile (P99) response duration in seconds.
Container Restarts: The aggregate rate of change in container restart counts over a 15-minute window.
Annotations & Ground Truth
The dataset is fully annotated with precise temporal intervals. We distinguish between two types of system events:
Anomalous Periods (Anomalies): Realistic faults such as gradual memory leakages, CPU synchronization stutters, and network latency jitter. These are labeled as is_anomaly = 1.
Benign Events (Noise): Normal but intensive system activities, such as Prometheus metric scraping and log rotation, which often trigger false positives in traditional detectors. These are labeled as is_anomaly = 0.
Data Quality & Artifacts
Researchers should note that the dataset contains certain measurement noise and monitoring artifacts. Specifically, occasional negative values in CPU utilization may appear; these are not physical system events but arise from counter resets or differential calculation errors (e.g., Prometheus rate() artifacts) during high-load periods.
File Structure
The AIClusterKPI (ACK) dataset is partitioned into training and testing subsets to facilitate reproducible machine learning experiments. The repository consists of the following four files:
groundtruth.csv: The master annotation file containing precise start and end timestamps for all recorded events. It includes metadata for each entry, such as the target node, event type, and a detailed description distinguishing between real anomalies and benign system noise.
train1.csv & train2.csv: These files comprise the primary training set, featuring 31-dimensional multivariate KPI traces. They provide the long-term operational baseline required for models to learn the “normal” behavioral patterns of a cloud-native AI cluster under varying workloads.
test.csv: T
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.