One half-split rotation plane from a token-major f32 projection to a
head-major packed-f16 attention operand. Grid (sequence/16, heads, 1),
block 256: each group of 16 threads owns one position, with each thread
owning adjacent components in both 32-wide halves. Query
and key use the same program with different projection offsets, so K can
have fewer heads without materializing repeated K.
1384def attentionHeadLayoutSM86RotatePlaneAdmitted =
1385 (lambda unrestricted seq : Nat . (lambda unrestricted heads : Nat .
1386 (lambda unrestricted projectionWidth : Nat . (lambda unrestricted planeOffset : Nat .
1387 (naturalAnd (naturalNonzero seq)
1388 (naturalAnd (naturalEqual (naturalModuloUnchecked seq 16) 0)
1389 (naturalAnd (naturalNonzero heads)
1390 (naturalAnd (naturalLessOrEqual (naturalAdd planeOffset (naturalMultiply heads 64)) projectionWidth)
1391 (naturalAnd (naturalLessOrEqual (naturalMultiply seq (naturalMultiply projectionWidth 4)) 4294967295)
1392 (naturalLessOrEqual (naturalMultiply heads (naturalMultiply seq 128)) 4294967295))))))))))The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.