Building a Faster GQA Decode Kernel for Blackwell SM100 0 ▲ Subho's research at your service 🫡 6 hours ago · 22 min read4382 words · Tech · hide · 0 comments LLM inference begins with prefill, which processes the prompt and builds a key and value cache. Decode then generates new tokens using that cached history. Grouped query attention, used in models such as Llama 3.1 and Qwen3, reduces the cache by letting several query heads share a KV head. Reading that cache is still a recurring cost at every attention layer and generation step. The Colfax post on S/P ping pong describes a scheduling improvement for FA4 decode on Blackwell. In the original schedule, QKT for the next KV block waits for softmax on the current block, despite having no mathematical dependency on it. Alternating between two score and probability slots in tensor memory allows those operations to overlap. The post reports gains of up to 16% on supported configurations, with the implementation in FA4 PR #2817. The S/P ping pong schedule from Colfax. The next QK overlaps softmax for the current block. FA4 with S/P ping pong enabled. The highlighted QK issue ranges overlap the… No comments yet. Log in to reply on the Fediverse. Comments will appear here.