Source/Packages

Realization.Nvidia.SM86.SoftmaxSM86

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/SoftmaxSM86.alpha

2,114 lines259 declarations74.8 KiBSHA-256 5aa09b49868d

def · lines 913–949

softmaxSM86MaskOne1024

Full file
coppelius D24: causal softmax over 1024-wide rows (attention scores at context 1024), 256 threads x FOUR consecutive elements each, one block per row, the row index taken in full from CTAID.X (the 256-wide variant masks from the low eight bits). Registers: R0 tid, R1 row, R5 = 4, R17 = tid*4 (first column), R16 = row*1024 + R17 (element index), R2:R3 input address, R4/R12/R13/R14 the four scores, R11 = row - column0, R9 scratch, R6 accumulator, R7/R8 reduction scratch, R10:R11 output address.

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.