Source/Packages

Realization.Nvidia.SM86.AttentionHeadLayoutSM86

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/AttentionHeadLayoutSM86.alpha

1,710 lines162 declarations85.3 KiBSHA-256 ad8cef82929b

family · lines 38–43

AttentionHeadLayoutSM86Kind

Full file
coppelius S5 rung 4b (D24): the ATTENTION HEAD LAYOUT kernels at coppelius's geometry — sequence 1024, 8 heads of 64 (half width 32), model width 512, so the fused QKV projection is a 1024 × 1536 fp32 row-major matrix whose row is [Q heads 0..7 | K heads 0..7 | V heads 0..7], each head 64 wide. The tensor-core GEMM (Realization.Nvidia.SM86.GEMM.HMMA.*) computes C = A · Bᵀ with A (M × K) and B (N × K) both fp16 row-major. So per head h: scores_h = Q_h · K_hᵀ A = Q_h (1024 × 64), B = K_h (1024 × 64) out_h = P_h · V_h A = P_h (1024 × 1024), B = V_hᵀ (64 × 1024) and the layout kernels here produce exactly those operands: QK ROTARY: for every position p, head h and component c < 32 the HF half-split rotation (Llama): q'[c] = q[c]·cos − q[c+32]·sin, q'[c+32] = q[c+32]·cos + q[c]·sin, with cos/sin tables of 1024 × 32 fp32 (angle p · theta^(−2c/64)); the same for K; both written fp16 head-major [head][position][64]. Grid 1024 (one block per position), block 128: thread = head (tid >> 4) × component pair (tid & 15 → c = 2g, 2g+1), four fp32 loads per plane, two packed half2 stores per plane. VALUE TRANSPOSE: V written fp16 as [head][component][position]. Grid 512 (one block per PAIR of positions, so a thread packs two adjacent positions into one half2 store), block 512: thread = head (tid >> 6) × component (tid & 63). Constant bank 0 (64-bit pointers): 0x160 first output (Q or V), 0x168 second output (K; unused by the value kernel), 0x170 QKV input, 0x178 cosine table, 0x180 sine table. No host fallback: the image is emitted by the sovereign encoder or the build refuses with a stable code.
38family AttentionHeadLayoutSM86Kind : Type 0
39constructor AttentionHeadLayoutQKRotary
40constructor AttentionHeadLayoutValueTranspose
41constructor AttentionHeadLayoutHeadToToken
42constructor AttentionHeadLayoutTokenToHead
43constructor AttentionHeadLayoutInverseRoPEQKVMerge

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.