coppelius S5 rung 4b (D24): the ATTENTION HEAD LAYOUT kernels at coppelius's
geometry — sequence 1024, 8 heads of 64 (half width 32), model width 512,
so the fused QKV projection is a 1024 × 1536 fp32 row-major matrix whose
row is [Q heads 0..7 | K heads 0..7 | V heads 0..7], each head 64 wide.
The tensor-core GEMM (Realization.Nvidia.SM86.GEMM.HMMA.*) computes C = A · Bᵀ
with A (M × K) and B (N × K) both fp16 row-major. So per head h:
scores_h = Q_h · K_hᵀ A = Q_h (1024 × 64), B = K_h (1024 × 64)
out_h = P_h · V_h A = P_h (1024 × 1024), B = V_hᵀ (64 × 1024)
and the layout kernels here produce exactly those operands:
QK ROTARY: for every position p, head h and component c < 32 the HF
half-split rotation (Llama): q'[c] = q[c]·cos − q[c+32]·sin,
q'[c+32] = q[c+32]·cos + q[c]·sin, with cos/sin tables of 1024 × 32 fp32
(angle p · theta^(−2c/64)); the same for K; both written fp16 head-major
[head][position][64]. Grid 1024 (one block per position), block 128:
thread = head (tid >> 4) × component pair (tid & 15 → c = 2g, 2g+1), four
fp32 loads per plane, two packed half2 stores per plane.
VALUE TRANSPOSE: V written fp16 as [head][component][position]. Grid 512
(one block per PAIR of positions, so a thread packs two adjacent
positions into one half2 store), block 512: thread = head (tid >> 6) ×
component (tid & 63).
Constant bank 0 (64-bit pointers): 0x160 first output (Q or V), 0x168 second
output (K; unused by the value kernel), 0x170 QKV input, 0x178 cosine table,
0x180 sine table.
No host fallback: the image is emitted by the sovereign encoder or the
build refuses with a stable code.
38family AttentionHeadLayoutSM86Kind : Type 0
39constructor AttentionHeadLayoutQKRotary
40constructor AttentionHeadLayoutValueTranspose
41constructor AttentionHeadLayoutHeadToToken
42constructor AttentionHeadLayoutTokenToHead
43constructor AttentionHeadLayoutInverseRoPEQKVMergeThe compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.