1 day ago · 16 min read3204 words · Tech · hide · 0 comments

Using LLMs to find bugs and vulnerabilities in software is no longer new and cool, and my 2025 post about using specialized tools to look for vulnerabilities in open-source codebases seems like a distant fever dream. The new cool thing is creating static or deterministic checks (gates, oracles, whatever you want to call them) to verify that an issue discovered by an LLM is really a vulnerability or not. This post outlines all of the lessons I’ve learnt while attempting to create these deterministic checks, specifically for vulnerabilities which can be discovered using memory sanitizers - ASan, MSan, UBSan, TSan, and LSan. The overall concept is that LLMs happily find bugs, but many of those bugs are either not real, or not reachable – or just not a vulnerability. Asking a secondary agent to prove whether the issue is real or not doesn’t work: agents cheat, they flip-flop regardless of facts, and are too happy to make you happy by saying something is a “proved” vulnerability – even if…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.