Source/Packages

Realization.Nvidia.SM86.AttentionHeadLayoutSM86

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/AttentionHeadLayoutSM86.alpha

1,710 lines162 declarations85.3 KiBSHA-256 ad8cef82929b

def · lines 1569–1577

attentionHeadLayoutSM86TransposeValueAdmitted

Full file
Transpose one token-major f32 value plane to head-major [head][component] [position] packed BF16/f16. Grid (sequence/2, heads, 1), block 64. Each thread packs two adjacent positions of one component into a 32-bit store.
1569def attentionHeadLayoutSM86TransposeValueAdmitted =
1570  (lambda unrestricted seq : Nat . (lambda unrestricted heads : Nat .
1571    (lambda unrestricted projectionWidth : Nat . (lambda unrestricted planeOffset : Nat .
1572      (naturalAnd (naturalNonzero seq)
1573        (naturalAnd (naturalEqual (naturalModuloUnchecked seq 2) 0)
1574          (naturalAnd (naturalNonzero heads)
1575            (naturalAnd (naturalLessOrEqual (naturalAdd planeOffset (naturalMultiply heads 64)) projectionWidth)
1576              (naturalAnd (naturalLessOrEqual (naturalMultiply seq (naturalMultiply projectionWidth 4)) 4294967295)
1577                (naturalLessOrEqual (naturalMultiply heads (naturalMultiply seq 128)) 4294967295))))))))))

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.