Maple-Preview 0 ▲ Tao of Mac 1 hour ago · Tech · hide · 0 comments This intrigues me quite a bit, since I’ve been looking into ternary models myself. It isn’t hard (at all) to beat Gemma 4 on quality, or to push tokens quickly through consumer hardware if you accept the quality trade-offs, but there may be a sweet spot at the intersection of quantization, model size and–more importantly–memory bandwidth where local models become genuinely useful without what is now massively expensive prosumer hardware. As usual, I have quibbles with the benchmarks–Qwen’s MoE models, especially with MTP, are hard to beat at this scale, and Maple’s own table has Qwen3.5 35B-A3B ahead overall. I’m taking the 218 tokens/s and SOTA framing with a large pinch of salt, but this is another useful data point on the number of optimization techniques people are trying. No comments yet. Log in to reply on the Fediverse. Comments will appear here.