1 hour ago · 6 min read1174 words · Tech · hide · 0 comments

A while back, I refactored REDbot to separate out a HTTP message linting library. httplint checks about 325 aspects of HTTP messages, covering the syntax and semantics of almost 200 HTTP header fields. Then, I got curious: what happens when you point it at Common Crawl? Its availability on AWS brings the opportunity to do a large-scale survey of how HTTP is used and abused. With the help of Claude Code, the result is cc-lint. Pointed at over 120 million responses from nearly 50,000 sites, the resulting report is a lot to digest. For me, the most interesting aspects are reflections of how we design new protocol elements – and often, how they come up short. Collectively, what I observe below is in some ways very obvious, but also we keep on forgetting it when we design new protocol elements – i.e., “hope-based standardisation.” So, this post serves as a reminder or wake-up call for protocol designers (depending on your perspective). It covers the tool and those observations; followups…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.