1 hour ago · Tech · hide · 0 comments

About a month ago I discovered that the scripts that classify pages as either “text” or “comics” on kwakk.info‘s section for “text pages from comics” (*phew*) didn’t like pages like the above. Which meant that all letters pages that had illustrations and stuff were categorised as comics and left out of the search index. So, for instance, none of the Johnny DC pages were included. Which is a serious mishap as you can imagine! Even worse — Usagi Yojimbo letters pages like the above were also excluded. Well, a new classification script was put into production, and it chewed through nine million comics pages and spat out a new list. Those pages were then put through the OCR and indexing processes. The computer worked at this for three weeks. The results seem promising — there’s now more pages included, of course, but it might well miss some other pages now. I mean, classifying pages into “text” and “not text” is a probabilistic enterprise, so it’s going to make mistakes: Some previously…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.