Overview
Historical Polish is well documented, yet the machine-readable annotated record covers only about a million words for the period this paper addresses; the remainder is accessible only through optical character recognition (OCR) of variable quality. To close this gap, Szymon Kocur introduces Wieszcz-XIX, a corpus of 6.75 billion tokens (approximately 3.1 billion words) drawn from 294,369 documents—mostly periodical issues—published in Polish between 1800 and 1918. The corpus is assembled from Wolne Lektury and the Internet Archive using a pipeline that filters, deduplicates, audits for post-1918 leakage, and splits at the document level. It is more than three orders of magnitude larger than the annotated corpus covering the same period.
Corpus Construction and Quality Audit
The authors quantify the corpus's defects, including:
- Recognition corruption, measured against a false-positive floor
- Near-identical duplication, which is removed
- Post-1918 leakage, excluded from the training corpus itself down to a known residue of 0.04–0.38% of its bytes, found in the transcribed source
As a result, the published corpus is the trained one document for document. On a hand-corrected sample, the character error rate is 0.68% where the text is legible, while 45% of sampled passages cannot be corrected at all.
Temporally Bounded Language Models
Using this corpus, the study trains a ladder of decoder-only models from scratch, ranging from 47M to 349M parameters, and measures their temporal boundedness. Against two modern Polish base models—one substantially larger—the 349M model shows a crossover, as does the 107M model against the comparator of its size. Key findings include:
- Post-1918 vocabulary costs the trained models roughly 3.1 bits per byte more than period vocabulary—a gap the comparators do not exhibit.
- Period vocabulary costs the trained models fewer bits than it costs the comparators.
- Shown period text, the models retain its spelling, whereas the comparators do so only partly.
- Adding parameters yields about twice the gain of a second pass over the data.
Resources and Ethical Considerations
The corpus, code, and model weights are publicly released. A content warning accompanies the release: the models reproduce period prejudice, including antisemitic statements.
Paper details: 35 pages, 2 figures, 11 tables. Submitted 6 October 2026. Subjects: Computation and Language (cs.CL); Digital Libraries (cs.DL). ACM classes: I.2.7; H.3.7. arXiv:2610.10592 [cs.CL]. Resources: Corpus · Code · Weights
via ArXiv CL+LG
