1 hour ago · 10 min read2060 words · Tech · hide · 0 comments

Last week DeepSeek-V4.1-Flash went live, and as it happens whenever DeepSeek releases a new model, the AI sphere went crazy. Going from v4 to v4.1 may not seem like a major release, but according to many experts in the field, the changes in the architecture could have easily justified this model being DeepSeek-V5.With this, I couldn’t miss the opportunity to extend the series of posts that I started with Kimi K3, the post-training piece and the Flash pair, and dedicate a post explaining what the changes in the architecture of this DeepSeek model means for the field. Once more, this model prioritises performance improvements and is designed in a way that roughly a third of this model doesn’t need to be hosted in the GPU.Some of the tricks introduced in this new DeepSeek architecture could already be seen in the Qwen3.8-Flash-Next architecture I described in this post, which is interesting because, and this is going to become my most repeated sentence this year, “it seems like we are…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.