Beyond Bigger MoE: How Kimi K3 Scales Context, Depth, and Agents 0 ▲ Andrey Lukyanenko 1 hour ago · 7 min read1306 words · Tech · hide · 0 comments Beyond Bigger MoE: How Kimi K3 Scales Context, Depth, and Agents Paper Code Project Recent frontier-model development has increasingly emphasized reinforcement learning and test-time computation: make a strong pretrained model reason longer, use more tools, and execute increasingly complex trajectories. Kimi K3 argues that post-training alone is not enough. Moonshot AI scales both sides simultaneously: the pretrained foundation grows to 2.8T total parameters with 104B activated per token, while post-training expands to 1M-token agentic trajectories, multiple reasoning-effort levels, coding, knowledge work, and general tool use. Kimi K3 organizes scaling around three kinds of information flow: Across sequence length, it interleaves three efficient Kimi Delta Attention layers with one global Gated MLA layer. Across depth, Attention Residuals let layers selectively retrieve representations from earlier blocks rather than receiving all previous computation through a single accumulated… No comments yet. Log in to reply on the Fediverse. Comments will appear here.