Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Half-Truth: A Partially Fake Audio Detection Dataset (HAD)

Several promising datasets have been developed to advance the field of fake audio detection. However, these previous datasets have failed to address a critical scenario: the presence of an attacker who covertly inserts small fabricated audio clips into authentic speech recordings. This situation

👤
CreatorYi, Jiangyan
📅
Published2023-12-14
🔗
DOI10.48550/arXiv.2104.03617
📊
Downloads5,557
⚖️
Licensecc-by-4.0
File Size7.5 GB
Data TypeDataset
Published2023
Licensecc-by-4.0
Total Views3,986
Total Downloads5,557

Several promising datasets have been developed to advance the field of fake audio detection. However, these previous datasets have failed to address a critical scenario: the presence of an attacker who covertly inserts small fabricated audio clips into authentic speech recordings. This situation poses a significant security threat because differentiating these small fake segments from the overall speech utterance is an exceptionally challenging task. In response to this challenge, we introduce a groundbreaking dataset designed for the detection of partial audio falsifications, which we term Half-Truth Audio Detection (HAD). The partially manipulated audio samples contained within the HAD dataset involve minimal alterations, typically limited to modifying a few words within an utterance. These altered audio segments are created using state-of-the-art speech synthesis technology.This dataset not only empowers the identification of counterfeit utterances but also enables the pinpointing of manipulated regions within a speech recording.

When you use this dataset, please cite us:

Jiangyan Yi, Ye Bai, Jianhua Tao, Haoxin Ma, Zhengkun Tian, Chenglong Wang, Tao Wang, Ruibo Fu:Half-Truth: A Partially Fake Audio Detection Dataset. Interspeech 2021: 1654-1658

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Half-Truth: A Partially Fake Audio Detection Dataset (HAD) (Full Dataset)7.5 GB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Yi, Jiangyan (2023). Half-Truth: A Partially Fake Audio Detection Dataset (HAD). https://doi.org/10.48550/arXiv.2104.03617