Forschende arbeiten gemeinsam auf einem Universitätscampus

Projekt

GutenbergKG: The Knowledge Press

GutenbergKG is a universal ingestion engine for digitized text corpora. It downloads public-domain works from Project Gutenberg and the Internet Archive, converts them to structured Markdown, and indexes them as queryable knowledge graphs via DocKG and KGRAG. The current corpus spans 253 works across 21 genres — 1,344…

GutenbergKG is a universal ingestion engine for digitized text corpora. It downloads public-domain works from Project Gutenberg and the Internet Archive, converts them to structured Markdown, and indexes them as queryable knowledge graphs via DocKG and KGRAG. The current corpus spans 253 works across 21 genres — 1,344,428 nodes, 5,280,014 edges — queryable as independent genre corpora or as a unified knowledge graph.