Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Constrainify Case Study: Data Quality Reports

This deposit contains the machine-readable data quality reports produced in the case study evaluating Constrainify, a web-based, self-service system for defining and performing data quality analysis on object description data.

👤
CreatorMatoni, Markus
📅
Published2026-05-31
🔗
DOI10.5281/zenodo.20477578
📊
Downloads164
⚖️
Licensecc-by-4.0
File Size31.6 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views58
Total Downloads164

This deposit contains the machine-readable data quality reports produced in the case study evaluating Constrainify, a web-based, self-service system for defining and performing data quality analysis on object description data.

The study was conducted as part of the AQinDa project. It operationalizes the abstract data quality requirements published by the German Digital Library (DDB) (see Anforderungen an die Lieferdaten) into an executable constraint set for LIDO data and executes this set on five real-world datasets delivered by DDB data providers. Each report enumerates, per constraint and per analyzed file, the detected quality incidents that serve as evidence of data quality problems.

Redaction and Pseudonymization

The reports are derived data published in redacted form to protect the underlying provider records.

  • Incident snippets. In the original reports, each detected incident carries a snippet field containing a raw excerpt of the analyzed record. Every such snippet has been replaced by the placeholder value "redacted", so that the concrete content of individual records is not disclosed.
  • File identifiers. The original file identifiers, which could be resolved back to specific provider records, have been replaced by sequential pseudonyms with a dataset prefix (e.g. D1_1.xml, D1_2.xml). Within a report, the same original file consistently maps to the same pseudonym.
  • Datasets. The five datasets are anonymized as D1–D5.

All remaining fields (constraint identifiers and names, compliance and incident counts, and execution metadata) are preserved unchanged, so that the analytical results remain fully interpretable.

Dataset

The five analyzed datasets are summarized below; record counts and volumes refer to the source datasets.

  • D1: Numismatics: 281 records; 11 MB; de, en.
  • D2: Numismatics: 6,955 records; 309 MB; de, en.
  • D3: Classical Antiquities: 10,565 records; 152 MB; de, en.
  • D4: Archaeology: 9,149 records; 343 MB; de, en.
  • D5: Architectural Photography: 46,438 records; 255 MB; de.

File Structure

Each dataset is provided as one gzip-compressed JSON file (report_D1_REDACTED.json.gz – report_D5_REDACTED.json.gz). Each decompressed file is a JSON object with a single key "result", whose value is an array of constraint&

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Constrainify Case Study: Data Quality Reports (Full Dataset)31.6 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Matoni, Markus (2026). Constrainify Case Study: Data Quality Reports. https://doi.org/10.5281/zenodo.20477578