Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Text Reuse in GEI-Digital Textbooks from Lower Saxony Detected with Ikhneutes

This file contains a corpus of historical German textbooks annotated with text reuse detected by n-gram shingling. It was created in the project “The Evolution of the T

👤
CreatorDombrowski, Fabian
📅
Published2026-04-10
🔗
DOI10.5281/zenodo.19570793
📊
Downloads10
⚖️
Licensecc-by-4.0
File Size9.5 GB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views111
Total Downloads10

This file contains a corpus of historical German textbooks annotated with text reuse detected by n-gram shingling. It was created in the project “The Evolution of the Textbook. Text reuse detection for the analysis of the influence of early textbook production in Lower Saxony on the development of the genre as a whole [TextbookEvolution]”, funded by the Lower Saxony Ministry for Science and Culture. Text reuse detection was carried out with Ikhneutes, a Rust-based pipeline and viewer for text reuse detection, schedule for release for spring/summer 2026. The corpus comprises 462 textbooks published between 1762 and 1918 in what today is Lower Saxony, Germany: mostly in Braunschweig, Göttingen, Hannover, Oldenburg, Osnabrück, and Wolfenbüttel. The subjects are geography, history, “Realien” (natural sciences), German primers, and reading books. The books were originally digitized by the Leibniz Institute for Educational Media | Georg Eckert Institute (GEI) for their digital collection of historical textbooks “GEI-Digital”; the underlying source data derive from the Projektkorpus “Schulbuch-Evolution” aus “GEI-Digital” – Annotierte Daten, that is, a selective data extraction from GEI-Digital whose TEI-XML sources were linguistically annotated by Frank Wiegand of the Zentrum für digitale Lexikographie der deutschen Sprache (ZDL), using DTA::CAB and distributed in DDC-Tabs format.

Corpus statistics:

  • Total books: 462
  • Total pages: 114,381
  • Total tokens: 50,529,431
  • Detected reuse instances (individual correlated passages): 106,491 (these have not yet been systematically evaluated for false positives)
  • Books with detected reuse: 460
  • Book pairs with reuse relations: 33,447

The file itself is an export of a SQLite 3 database generated by Ikhneutes. It is suitable for archival distribution, SQL-based inspection, and downstream processing. The database contains document records and metadata, token and lemma occurrences with positional information, and computed text-reuse annotations between documents. It was built fr

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Text Reuse in GEI-Digital Textbooks from Lower Saxony… (Full Dataset)9.5 GB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Dombrowski, Fabian (2026). Text Reuse in GEI-Digital Textbooks from Lower Saxony Detected with Ikhneutes. https://doi.org/10.5281/zenodo.19570793