Source/Packages

Realization.Nvidia.SM86.StreamingAttentionSM86

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/StreamingAttentionSM86.alpha

1,129 lines205 declarations66.1 KiBSHA-256 aefcbc3b00ca

def · lines 786–786

skK

Full file
The key launch (streamingAttentionKeySM86): a block of two warps per 32 keys of a K/V head, grid (T / 32, K/V heads, groups). A warp holds its 16 keys' K and V rows as A fragments and walks the 32-query sub-tiles from the last one down to its diagonal (the last, masked): S^T = K Q^T, dP^T = V dO^T, P^T = 2^(S^T c - L), dS^T = P^T (dP^T - D) with L and D the queries', dV += P^T dO and dK += dS^T Q. Operands: Q, K, V, dO ([head][T][64], half), Q^T, dO^T ([head][64][T], half), L and D ([head][T]); results dK ([head][T][64]) and dV^T ([head][64][T]), binary32. Parameters: Q, K, V, dO, Q^T, dO^T, L, D, dK, dV^T, then c and s.
786def skK = (lambda unrestricted kt : Nat . (naturalAdd 20 (naturalMultiply kt 4)))

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.