1 day ago · Tech · hide · 0 comments

In 1998 I needed to convert PDF files to text for a search engine called Alkaline, without requiring separately installed libraries. So I wrote a C++ parser. Alkaline is another story. I recovered the source from a backup dated December 26, 1998. It reads PDF 1.0 through 1.2, follows object references and page trees, decompresses content with bundled zlib, and prints metadata, page text and bookmark titles. There are my own string, vector and hash-table libraries in there too. The encryption code is a stub. This was a work in progress. I wasn’t reverse engineering an undocumented format. Adobe had published the specification in 1993, and my comments refer to the PDF Reference, right down to a page number for stream filters. This predates tagged PDF, introduced in 2001. Getting text meant interpreting page drawing commands, not following a tagged reading order. I was very impressed by how clever the format was. You start near the end, find the cross-reference table, and use its byte…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.