Source/Packages

Realization.Nvidia.SM86.LinearStepLoopSM86

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/LinearStepLoopSM86.alpha

178 lines33 declarations8.0 KiBSHA-256 059e5870f8c0

def · lines 35–35

llR11

Full file
The linear step with LOOPS: the same computation as Realization.Nvidia.SM86.LinearStepSM86 -- the same operations on the same operands in the same order, so the record is the specification's word for word -- in a program of constant length and constant register demand. Where the unrolled program keeps W, x, t, dW and W' in registers (848 of them for 16 x 16; the register file admits 38 x 1 .. 1 x 57), this one walks the outputs in an outer loop and the inputs in two inner loops, loading each element when it is used and storing each result when it is made, with 64-bit addresses from IMAD.WIDE (index x 4 + the parameter word's address). The loops close with ISETP.GT (an unsigned compare -- Checked.IntegerCompareProbe) and a backward BRA. prologue (tid, the lane guard, the parameter words; LinearStepSM86) R11 = -eta R12 = 0.5 R13 = loss = +0 R35 = 4 R14 = i = 0 R16 = e = 0 outer: R28 = y = +0 R15 = j = 0 first: W_e, x_j loaded; y = fma(W_e, x_j, y); e += 1; j += 1; again while j <= k-1 y_i stored; t_i loaded; d = y + (-t_i); loss = fma(d, d, loss); e -= k; j = 0 second: x_j, W_e loaded; dW = d x x_j; W' = fma(dW, -eta, W_e); both stored; e += 1; j += 1; again while j <= k-1 i += 1; again while i <= m-1 loss = 0.5 x loss, stored at the record's word m (i = m by then); exit Each backward branch's offset is its loop's length in bytes, negated, counted from the loop's own block (llLoop).
35def llR11 : Nat = 11

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.