Constrainify Case Study: Data Quality Reports
This deposit contains the machine-readable data quality reports produced in the case study evaluating Constrainify, a web-based, self-service system for defining and performing data quality analysis on object description data.
This deposit contains the machine-readable data quality reports produced in the case study evaluating Constrainify, a web-based, self-service system for defining and performing data quality analysis on object description data.
The study was conducted as part of the AQinDa project. It operationalizes the abstract data quality requirements published by the German Digital Library (DDB) (see Anforderungen an die Lieferdaten) into an executable constraint set for LIDO data and executes this set on five real-world datasets delivered by DDB data providers. Each report enumerates, per constraint and per analyzed file, the detected quality incidents that serve as evidence of data quality problems.
Redaction and Pseudonymization
The reports are derived data published in redacted form to protect the underlying provider records.
- Incident snippets. In the original reports, each detected incident carries a
snippetfield containing a raw excerpt of the analyzed record. Every such snippet has been replaced by the placeholder value"redacted", so that the concrete content of individual records is not disclosed. - File identifiers. The original file identifiers, which could be resolved back to specific provider records, have been replaced by sequential pseudonyms with a dataset prefix (e.g.
D1_1.xml,D1_2.xml). Within a report, the same original file consistently maps to the same pseudonym. - Datasets. The five datasets are anonymized as D1–D5.
All remaining fields (constraint identifiers and names, compliance and incident counts, and execution metadata) are preserved unchanged, so that the analytical results remain fully interpretable.
Dataset
The five analyzed datasets are summarized below; record counts and volumes refer to the source datasets.
- D1: Numismatics: 281 records; 11 MB; de, en.
- D2: Numismatics: 6,955 records; 309 MB; de, en.
- D3: Classical Antiquities: 10,565 records; 152 MB; de, en.
- D4: Archaeology: 9,149 records; 343 MB; de, en.
- D5: Architectural Photography: 46,438 records; 255 MB; de.
File Structure
Each dataset is provided as one gzip-compressed JSON file (report_D1_REDACTED.json.gz – report_D5_REDACTED.json.gz). Each decompressed file is a JSON object with a single key "result", whose value is an array of constraint&
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.