Wieszcz-XIX: 6.75B Token Corpus of Historical Polish Language
October 9, 2026
The Wieszcz-XIX corpus provides 6.75 billion tokens of Polish text from 1800 to 1918, extracted from various digital archives. The pipeline includes deduplication and leakage detection, limiting post-1918 data residue to under 0.38% of bytes.
HOW THIS AFFECTS YOU
●
builderYou can now train or fine-tune temporally bounded models on high-quality historical Polish.
●
researcherThis dataset allows for much more granular studies of linguistic evolution in Polish.