9 hours ago · Tech · hide · 0 comments

I ran five different coding agents at the same local model, on the same task, with the same frozen test suite — and then I counted why they failed. The answer wasn’t subtle. About 90% of the failures were harness problems. Only about 10% were the model. And here’s the part that should change how you spend your next dollar: throwing a bigger, less-quantized model at it fixed none of them.

No comments yet. Log in to reply on the Fediverse. Comments will appear here.