More Muddling with NDQ 0 ▲ Archaeology of the Mediterranean World 1 hour ago · 6 min read1173 words · Tech · hide · 0 comments I spent the last few days alternating between reading, thinking, and muddling around with NDQ using Claude Code. About 50% of my time involves cleaning up the data so that I can do more sophisticated analysis with this text. To get a sense for what I’ve been working on, check out the post on Tuesday. One of the oddities that I encountered with the first analysis was The Bellman clustered by itself with little overlap with any of the other contemporary magazines in the corpus (NDQ, the two NDSHS journals, DeLestry’s Western Magazine, and The Midland). Interrogating this further, I discovered that it represented a larger problem. Apparently the model I was running was only reading part of the text (the first 2000 or so words) to create the model. This was not acceptable and a good reminder to verify then trust. To address this, we devised a plan to break the entire 20 million word corpus into around 100k, 200 word segments and to perform a BERT analysis on all of these segments. As part… No comments yet. Log in to reply on the Fediverse. Comments will appear here.