UCSB netFound Pretraining Data Sampler
This is a sampler of pretraining data for netFound network foundation model. This data contains network packet headers in .arrow format and excludes payload or IP addresses, as per netFound preprocessing pipeline. This data is supposed to be used with netFound tokenizer (or any derivative tokeniz
This is a sampler of pretraining data for netFound network foundation model. This data contains network packet headers in .arrow format and excludes payload or IP addresses, as per netFound preprocessing pipeline. This data is supposed to be used with netFound tokenizer (or any derivative tokenizers).
Please, see https://github.com/SNL-UCSB/netFound for instructions on usage and location of the full pretraining dataset.
Data facts:
- Total collection time of the sampler: 2hrs.
- Total network flows in the sampler: 60,199,405
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.