Source/Packages

Realization.Nvidia.SM86.LinearStepSM86

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/LinearStepSM86.alpha

354 lines84 declarations21.5 KiBSHA-256 a13116730702

def · lines 118–118

lsBase

Full file
---- the program, for any shape ---- Registers, for outputs m and inputs k: R0 tid; R2:R3 &W; R4:R5 &x; R6:R7 &t; R8:R9 &out; R10 eta; then x (k), t (m), W (m x k), y (m), d (m), loss, 0.5, dW (m x k), -eta, W' (m x k), consecutively from R12 -- for 2 x 2: R12,R13 x; R14,R15 t; R16..R19 W; R20,R21 y; R22,R23 d; R24 loss; R25 0.5; R26..R29 dW; R30 -eta; R31..R34 W', the plan the RTX 3090 ran. The program is generated from the shape; every loop below is unrolled, in the semantic program's order (fma chains first element innermost).
118def lsBase : Nat = 12

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.