5 hours ago · Tech · hide · 0 comments

Over the last two articles we went through everything needed to run a model: an SDK that runs inference, and a management layer that downloads models, sizes them for your hardware and loads them. You can use the SDK directly from your Go code, but one of the most popular ways of consuming LLMs is through an HTTP API, and that’s what’s left: the server that lets your existing OpenAI client talk to all of it.

No comments yet. Log in to reply on the Fediverse. Comments will appear here.