coppelius S5 rung 4b (D24): the ATTENTION HEAD LAYOUT kernels at coppelius's
geometry — sequence 1024, 8 heads of 64 (half width 32), model width 512,
so the fused QKV projection is a 1024 × 1536 fp32 row-major matrix whose
row is [Q heads 0..7 | K heads 0..7 | V heads 0..7], each head 64 wide.
The tensor-core GEMM (Realization.Nvidia.SM86.GEMM.HMMA.*) computes C = A · Bᵀ
with A (M × K) and B (N × K) both fp16 row-major. So per head h:
scores_h = Q_h · K_hᵀ A = Q_h (1024 × 64), B = K_h (1024 × 64)
out_h = P_h · V_h A = P_h (1024 × 1024), B = V_hᵀ (64 × 1024)
and the layout kernels here produce exactly those operands:
QK ROTARY: for every position p, head h and component c < 32 the HF
half-split rotation (Llama): q'[c] = q[c]·cos − q[c+32]·sin,
q'[c+32] = q[c+32]·cos + q[c]·sin, with cos/sin tables of 1024 × 32 fp32
(angle p · theta^(−2c/64)); the same for K; both written fp16 head-major
[head][position][64]. Grid 1024 (one block per position), block 128:
thread = head (tid >> 4) × component pair (tid & 15 → c = 2g, 2g+1), four
fp32 loads per plane, two packed half2 stores per plane.
VALUE TRANSPOSE: V written fp16 as [head][component][position]. Grid 512
(one block per PAIR of positions, so a thread packs two adjacent
positions into one half2 store), block 512: thread = head (tid >> 6) ×
component (tid & 63).
Constant bank 0 (64-bit pointers): 0x160 first output (Q or V), 0x168 second
output (K; unused by the value kernel), 0x170 QKV input, 0x178 cosine table,
0x180 sine table.
No host fallback: the image is emitted by the sovereign encoder or the
build refuses with a stable code.
42constructor AttentionHeadLayoutTokenToHeadThe compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.