HyperFM250k_subset
HyperFM Dataset Overview This dataset contains **4,250 hyperspectral samples** from NASA PACE mission, curated for machine learning and computer vision research. It is designed to support tasks such as regression, segmentation, and representation learning using
HyperFM Dataset
Overview
This dataset contains **4,250 hyperspectral samples** from NASA PACE mission, curated for machine learning and computer vision research. It is designed to support tasks such as regression, segmentation, and representation learning using hyperspectral imaging (HSI) data.
A larger version of this dataset (**3TB+**) can be generated using the provided codebase (link below).
Directory Structure
.
├── hsi/
├── target/
├── train_list_2k.csv
├── valid_list_250.csv
├── test_list_2k.csv
Description
– **hsi/**
Contains hyperspectral data samples. Each file corresponds to a single observation. Dim: (96,96,291)
– **target/**
Contains labels or target values associated with each HSI sample. Four (4) Regression Tasks: *cld_xxx.npy [Dim: (96,96,4)]. One (1) Segmentation Task: *cldmask_xxx.npy [Dim: (96,96)]
– **train_list_2k.csv**
Training split with ~2,000 samples.
– **valid_list_250.csv**
Validation split with ~250 samples.
– **test_list_2k.csv**
Test split with ~2,000 samples.
Each CSV file defines the mapping between input samples and their corresponding targets.
Dataset Statistics
– Total samples: **4,250**
– Training set: ~2,000 samples
– Validation set: ~250 samples
– Test set: ~2,000 samples
Full Dataset Generation
A significantly larger dataset (**3TB+**) can be generated using the official codebase:
– GitHub repository: *https://github.com/umbc-sanjaylab/HyperFM*
The repository includes scripts for:
– Data preprocessing
– Hyperspectral cube generation
– Label construction
– Dataset scaling
Associated Publication
– Paper: Zahid Hassan Tushar, and Sanjay Purushotham, “HyperFM: A Efficient Hyperspectral Foundation Model with Spectral Grouping”, CVPR 2026 (findings).
Please cite the paper if you use this dataset in your research.
Usage Notes
– Ensure consistent preprocessing across splits when training models.
– Refer to the CSV files for correct pairing of inputs and targets.
– Large-scale experiments are recommended using the full generated dataset.
License
CC BY 4.0
Contact
For questions, issues, or collaboration:
– **Zahid Hassan Tushar**
Email: ztushar1@umbc.edu
– **Sanjay Purushotham**
Email: psanjay@umbc.edu
Acknowledgments
If you use this dataset, please acknowledge the authors and associated publication.
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.