The indexed-row scatter with a DETERMINISTIC reduction: destination row
ids[j] += source row j for every row j, as Realization.Nvidia.SM86.
IndexedRowScatterSM86 computes it with RED.E.ADD.F32 -- whose additions
into a row that several source rows share land in whatever order the
blocks run -- but in one order, fixed by the input: each destination row
belongs to the block of the FIRST source row that names it, which adds
every source row naming it, in ascending order, to the row's old words.
The other blocks exit. So each destination word is written by one thread
once, no two blocks touch the same row, and the result is
dest[v][c] = (((dest[v][c] + src[j0][c]) + src[j1][c]) + ...)
for the rows j0 < j1 < ... that name v -- the same bits every run
(Learning.Checked.IndexedRowScatter; SM86.IndexedRowScatterCheck decides
it on the block model, block by block in either order).
Same ABI as the atomic form (destination, source and index pointers at
c[0x160], c[0x168], c[0x170]); one block per source row, one thread per
component.
R0 = t (the block) R1 = c (the thread) R5 = 4 R12 = j = 0
R11 = v = ids[t] R25 = t + 1
again: u = ids[j]; exit when u = v and j != t; j += 1; again while j != t + 1
j = t; R20 = dest[v][c]
again: u = ids[j]; when u = v: R20 += src[j][c]; j += 1; again while j <= rows - 1
dest[v][c] = R20; exit
Equality is a nonzero exclusive or (ISETP.GT 0, unsigned); every loop's
backward branch is its own length (LinearStepLoopSM86.llLoop). The scan
takes rows 0 .. t, never none, so no branch jumps forward: only the
backward form's encoding is proven on the card (a forward jump with the
backward form's sign bits is a fault the thread model refuses).
44def irsT : Nat = 0The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.