MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators
Datasets Used This repository contains code, benchmark scripts, and experimental results associated with the paper: MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators</stro
Datasets Used
This repository contains code, benchmark scripts, and experimental results associated with the paper:
MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators
The experiments were conducted using the following publicly available datasets:
VoxCeleb1
Citation: Nagrani et al. (2017)
VoxCeleb1 is a large-scale audiovisual speaker recognition dataset containing speech recordings collected from interview videos. In this work, VoxCeleb1 was used to evaluate representation fidelity and downstream classification performance through a gender classification task.
Dataset URL:
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/
SPIRA
Citation: Casanova et al. (2021)
SPIRA is a respiratory health dataset composed of speech recordings collected for the assessment of respiratory insufficiency and COVID-19 related symptoms. In this work, SPIRA was used to evaluate the proposed MFCCT frontend in a clinical respiratory classification setting.
Dataset URL:
https://github.com/SPIRA-Project
LibriSpeech
Citation: Panayotov et al. (2015)
LibriSpeech is a corpus of read English speech derived from public-domain audiobooks. In this work, LibriSpeech samples were used as benchmark inputs for latency and energy measurements across multiple hardware platforms.
Dataset URL:
https://www.openslr.org/12
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.