Source/Packages

Realization.Nvidia.SM86.AttentionHeadLayoutSM86

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/AttentionHeadLayoutSM86.alpha

1,710 lines162 declarations85.3 KiBSHA-256 ad8cef82929b

def · lines 1384–1392

attentionHeadLayoutSM86RotatePlaneAdmitted

Full file
One half-split rotation plane from a token-major f32 projection to a head-major packed-f16 attention operand. Grid (sequence/16, heads, 1), block 256: each group of 16 threads owns one position, with each thread owning adjacent components in both 32-wide halves. Query and key use the same program with different projection offsets, so K can have fewer heads without materializing repeated K.
1384def attentionHeadLayoutSM86RotatePlaneAdmitted =
1385  (lambda unrestricted seq : Nat . (lambda unrestricted heads : Nat .
1386    (lambda unrestricted projectionWidth : Nat . (lambda unrestricted planeOffset : Nat .
1387      (naturalAnd (naturalNonzero seq)
1388        (naturalAnd (naturalEqual (naturalModuloUnchecked seq 16) 0)
1389        (naturalAnd (naturalNonzero heads)
1390          (naturalAnd (naturalLessOrEqual (naturalAdd planeOffset (naturalMultiply heads 64)) projectionWidth)
1391            (naturalAnd (naturalLessOrEqual (naturalMultiply seq (naturalMultiply projectionWidth 4)) 4294967295)
1392              (naturalLessOrEqual (naturalMultiply heads (naturalMultiply seq 128)) 4294967295))))))))))

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.