The SM86 realization of Learning.Checked.LinearStep for outputs 2, inputs
2: one thread of one block does the whole step in the semantic program's
exact operation order (fused multiply-adds first element innermost,
subtraction as addition of a negation, the loss as 0.5 x a fused sum of
squares, the update as fma(dW, -eta, W)); every other lane exits at once.
Preconditions (Realization contract):
outputs = 2, inputs = 2 (the register plan is for this shape);
one launch of one block, block width 32 (a single warp);
the parameter block carries, at the offsets below, the 64-bit addresses
of W (row-major, 16 bytes), x (8 bytes), t (8 bytes) and the 44-byte
output record (y 8, loss 4, dW 16, W' 16), and eta at 0x190;
the arena regions those addresses name are disjoint and 4-aligned.
Numerical contract: identical to the semantic program's (fused, ordered);
reproducibility: bitwise on any SM86 device for this realization, which
uses no reduction across lanes and no approximate unit.
29def linearStepSM86Identity : Bytes = b"linear-step-sm86-outputs2-inputs2-v1"The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.