Evaluating OCR on the Community Memory Corpus 0 ▲ ℤ→ℤ 1 hour ago · 21 min read4151 words · Tech · hide · 0 comments The Community Memory corpus consists of approximately 5,500 scanned printouts of bulletin board-style posts made by Berkeley residents in the 1980s. Since residents posted messages through publicly accessible terminals, the postings reflect a potentially distinct and broader population than other computer-facilitated networks of the time (e.g. Usenet, The WELL). In advance of a potential project to translate the scans into an indexed, machine-readable form, we evaluate the optical character recognition (OCR) performance and economy of five models (Document AI, Gemini 3.7 and 3.8, Mistral OCR, and Qwen3) on a sample of the corpus. We find that current generation general purpose LLMs perform better than specialized OCR engines and while Qwen3 and Gemini 3.7 have similar price/performance ratios, Gemini 3.8 is surprisingly a degradation to Gemini 3.7. We conclude that research tasks using the OCR output will need to be robust to at least four character errors per line or one word error… No comments yet. Log in to reply on the Fediverse. Comments will appear here.