The linear step with LOOPS: the same computation as
Realization.Nvidia.SM86.LinearStepSM86 -- the same operations on the same
operands in the same order, so the record is the specification's word for
word -- in a program of constant length and constant register demand.
Where the unrolled program keeps W, x, t, dW and W' in registers (848 of
them for 16 x 16; the register file admits 38 x 1 .. 1 x 57), this one
walks the outputs in an outer loop and the inputs in two inner loops,
loading each element when it is used and storing each result when it is
made, with 64-bit addresses from IMAD.WIDE (index x 4 + the parameter
word's address). The loops close with ISETP.GT (an unsigned compare --
Checked.IntegerCompareProbe) and a backward BRA.
prologue (tid, the lane guard, the parameter words; LinearStepSM86)
R11 = -eta R12 = 0.5 R13 = loss = +0 R35 = 4 R14 = i = 0 R16 = e = 0
outer: R28 = y = +0 R15 = j = 0
first: W_e, x_j loaded; y = fma(W_e, x_j, y); e += 1; j += 1; again while j <= k-1
y_i stored; t_i loaded; d = y + (-t_i); loss = fma(d, d, loss); e -= k; j = 0
second: x_j, W_e loaded; dW = d x x_j; W' = fma(dW, -eta, W_e); both stored;
e += 1; j += 1; again while j <= k-1
i += 1; again while i <= m-1
loss = 0.5 x loss, stored at the record's word m (i = m by then); exit
Each backward branch's offset is its loop's length in bytes, negated,
counted from the loop's own block (llLoop).
35def llR11 : Nat = 11The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.