1 day ago · Tech · hide · 0 comments

The Thales project is a private AI inference cluster built on two Dell PowerEdge R650 servers, each equipped with an NVIDIA T4 GPU. The goal is to provide low‑latency, high‑throughput inference endpoints for a variety of models—fast, deep, and embedding—behind a LiteLLM router. This post summarizes the current status of the build, the steps completed, and the next milestones.

No comments yet. Log in to reply on the Fediverse. Comments will appear here.