22,580: GPT-2 to Kimi K3, explained 0 ▲ Bogdan Buduroiu 2 hours ago · Tech · hide · 0 comments MLA, KDA, MoE, NoPE all attack either bytes transferred or bytes resident. MLA shrinks KV per token, KDA replaces a O(n)O(n)O(n) cache with an O(1)O(1)O(1) state. No comments yet. Log in to reply on the Fediverse. Comments will appear here.