1 hour ago · Tech · hide · 0 comments

This is the kind of inference optimisation work I have been doing in go-pherence and my personal llama.cpp fork, but at Cloudflare scale. I have been trying to squeeze more useful models into the hardware I already own; they are trying to squeeze more requests into a global GPU fleet, and many of the trade-offs are the same. I expect “mid-tier” open-weight models to become much more popular throughout the rest of the year as optimisations like these make them cheaper to serve, hopefully putting a ceiling on OpenAI and Anthropic pricing and forcing some competition. I would like this to help burst the AI bubble, but OpenAI and Anthropic’s fantasy pricing is hardly the only thing keeping it aloft.

No comments yet. Log in to reply on the Fediverse. Comments will appear here.