4 hours ago · Tech · hide · 0 comments

Qwen3.5 9B through Ollama (GGUF and MLX), llama-bench, mlx-lm and rapid-mlx on an M3 Ultra. Measured tokens per second, memory, and what MTP changes.

No comments yet. Log in to reply on the Fediverse. Comments will appear here.