Self-Hosted AI: Qwen 3.8 on a Rented Blackwell at 150 Tokens per Second 0 ▲ Larvitz Blog 1 hour ago · 15 min read3067 words · Tech · hide · 0 comments For a long time, “self-hosted AI” meant one of two things to me. Either a small model on a laptop, slow enough to make coffee between answers, or a lot of money spent on a GPU that then spends most of its life idling under my desk. Somewhere in the last few months, that changed. Open-weight models got good. Not “good for an open model”. Good. My current setup: Qwen 3.8 27B on a rented NVIDIA RTX PRO 6000 Blackwell at RunPod, started with one command, answering at around 150 tokens per second, and torn down again when I am done. The screenshot above is from a running pod. It wrote a haiku about the GPU shortage while occupying one of the GPUs in question. Table of Contents Table of Contents Good Enough Is Now Really Good The Stack Why This Particular Card The Silent Trap: Mixed-Precision Checkpoints One Command How Fast Is Fast? Data Sovereignty: The Part I Care About Most Where the Sovereignty Ends What It Costs The Forgotten Pod Problem Frontier Architect, Open-Weight Builder What I… No comments yet. Log in to reply on the Fediverse. Comments will appear here.