Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Misinformation Span Detection – EI22 & BOL4Y

BOL4Y & EI22 Description of each data provided in the paper.   Contents The repository includes the following data files: - dump.csv — Parsed and normalized fact-check metadata extracted from AosFatos HTML pages

👤
CreatorMatos, Breno
📅
Published2026-03-18
🔗
DOI10.5281/zenodo.19097541
📊
Downloads384
⚖️
Licensecc-by-4.0
File Size116.0 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views129
Total Downloads384

BOL4Y & EI22

Description of each data provided in the paper.
 

Contents

The repository includes the following data files:

– dump.csv — Parsed and normalized fact-check metadata extracted from AosFatos HTML pages
– aos_fatos_pages.zip — Raw HTML pages crawled from AosFatos (BOL4Y)
– escriba-csv.zip — Transcripts from escriba
– bol4y.csv.gz — Main BOL4Y dataset (compressed CSV)
– ei22.csv — Main EI22 dataset (CSV)
– bol4y-transcripts.zip — Transcripts associated with BOL4Y
– ei22-transcripts.csv — Transcripts associated with EI22
 

Data

dump.csv

This file contains parsed fact-checking metadata from AosFatos, used to build BOL4Y.
Original data is in Brazilian Portuguese; column names were translated to English for usability.
 
 
Column nameDescription
titleTitle of the fact check
datePublication date of the fact check
aos_fatos_linkURL of the original AosFatos page
fact_checkFact-checking paragraph
topic_ptTopic(s) of the false claim (Portuguese)
sourceOrigin of the claim (e.g., livestream, speech)
source_urlsURLs to the original content being fact-checked
repetition_countNumber of times the same claim was repeated
year_days_pairYear, month, and day of each claim occurrence. Includes repeated claims
pageCrawl page index where the fact check was found
fact_check_idUnique identifier for the fact check (shared with BOL4Y)

BOL4Y (bol4y.csv.gz)

CSV was compressed for easier storage. In it, you will find the following fields:
 
Column nameDescription
fact_check_idIdentifier matching dump.csv
fileSource file for the transcript 
transcription_sourceTranscription system (escriba or whisper)
transcription_indexSegment ID(s) associated with each segment. For some fact-checked segmentes, you’ll see more than one id listed in this field, following the concatenation approach discussed in the paper
transcription_textClaim o

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Misinformation Span Detection – EI22 & BOL4Y (Full Dataset)116.0 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Matos, Breno (2026). Misinformation Span Detection – EI22 & BOL4Y. https://doi.org/10.5281/zenodo.19097541