1 day ago · Tech · hide · 0 comments

How I run a coding LLM locally on a single RTX 3090 Ti — the model journey from Claude-distilled to a dense 27B coder, why MoE beat dense for speed, why I stayed on llama.cpp over vLLM on Ampere, and the quant type and sampler flag that quietly bit me.

No comments yet. Log in to reply on the Fediverse. Comments will appear here.