layer l's heads' attention as one streaming launch (Realization.Nvidia.
SM86.StreamingAttentionSM86): a block per 64 queries
of a head, the output written into the layer's merged half, each row's
base-2 log-sum-exp into its log-sum-exp (the backward's)
2024def cgStreamingAttention = (lambda unrestricted l : Nat .
2025 (cgLaunch
2026 cgImageAttentionForward
2027 (naturalDivideUnchecked cgSeq 64)
2028 cgHeads
2029 1
2030 (cgPointer
2031 (cgArgument 0)
2032 (cgQueriesOf l)
2033 (cgPointer
2034 (cgArgument 1)
2035 (cgKeysOf l)
2036 (cgPointer
2037 (cgArgument 2)
2038 (cgValuesTOf l)
2039 (cgPointer
2040 (cgArgument 3)
2041 (cgMergedHalfOf l)
2042 (cgPointer (cgArgument 4) (cgLogSumExpOf l) (cgSlot (cgArgument 5) cgStreamingScale cgNoSlots))))))))The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.