
BK
Benjamin Kiessling (ALMAnaCH)
· 1 min read
ResearcharXiv cs.CV
A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition
arXiv:2609.20064v1 Announce Type: new
Abstract: Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CRNN line recognizers, which are visually grounded and trainable on modest data. Lightweight recurrence-free recognizers promise the accuracy of larger models with the practical advantages of CRNNs, yet have not been comprehensively evaluated on historical writing. We adapt PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compare it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material. While PP-OCRv6 does not consistently outperform the baseline when trained from scratch, heterogeneous pretraining produces markedly better generalization. Comparisons with the Qwen3.5-based Medusa recognizer further show that fine-tuned PP-OCRv6 can outperform a large VLM tailored towards historical Latin-script HTR.
Original source
This story was published by arXiv cs.CV and written by Benjamin Kiessling (ALMAnaCH). SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


