You need to agree to share your contact information to access this dataset

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this dataset content.

Swedish Print OCR — training and evaluation code

The code behind the paper A Vision–Language OCR Model and an Open Corpus (Cullhed). This is the curated pipeline that produced the released models, benchmarks and corpus — exploratory and superseded experiment code is not included. Comments and docstrings are in Swedish; they document not just what each script does but why it is built the way it is.

Companion repositories

Layout, mapped to the paper

data/ — §2 Training data

Source acquisition (fetch_runeberg.py, refresh_runeberg_text.py, add_spraakbanken.py, extract_spraakbanken_regions.py, add_external.py), repair and deduplication (repair_runeberg.py, dedup_runeberg.py), per-work language and typeface classification (detect_language.py, classify_typeface.py, train_typeface_cnn.py, typsnitt_verk.py), integrity checks (scan_images.py, audit_dataset.py, kolla_traningspar.py), and the final seeded mixture build (build_dataset_final.py — 34% roman / 44% Swedish blackletter / 22% German blackletter, convention and century instructions).

cleaning/ — §2, page–text alignment screen

The three-step segmentation check used to remove pages whose ground truth demonstrably belongs to another page: segment (segmentera_pool.py), read each segment (las_koord.py), align against ground truth (radpar.py), then filter the training mixture on page similarity (rensa_train.py).

benchmark/ — §3 Evaluation

Work-disjoint, period-stratified principal benchmark (build_testset.py), the character ten-gram leakage check that removes same-text-different-title editions (find_text_leakage.py), and the external Wikisource benchmark (fetch_wikisource.py, build_extern_wikisource.py, build_extern_dataset.py) with its cross-system reference-outlier screen (rensa_extern.py).

training/ — §4 Model

The SLURM training launcher (train_chain.sbatch): CHURRO-3B base, frozen vision encoder, lr 8e-5 cosine, effective batch 128, bfloat16, sequence packing, on 4× GH200.

eval/ — §4.1 and §5

The single scoring program used for every system (eval_testset.py, eval_cer.py), repetition-loop detection (loopvakt.py) backing the uniform decoding repair, baseline runners under identical conditions (run_kraken.py, run_surya.py, run_tesseract.py, run_trocr.py, run_got_ocr2.py, run_calamari.py + calamari_lines.py), and results aggregation (sammanstall.py).

corpus/ — §6 Released corpus

The production pipeline over Litteraturbanken's facsimile collection: work listing and download (lista_verk.py, lb_api.py), volume OCR (ocra_volymer.py), page-count control and gap filling (granska_sidkarta.py, komplettera.py), repair of looping/truncated pages (loopplan.py, las_om_loopar.py), page-level licence classification for split-licence volumes (klassa_apparat.py), verification of page naming against Litteraturbanken's proofread e-text (hamta_etexter.py, kolla_justering.py), and corpus assembly and release filtering (bygg_textkorpus.py, slapp_ur_textkorpus.py).

Environment

Python 3.12 on aarch64 (NVIDIA GH200), trained with ms-swift; inference via vLLM. Paths in the scripts refer to the original project layout (dataset/, external/, work/ under a project root) — pair with the training-data repository's image shards, which preserve those relative paths.

Downloads last month
9