The SM86 realization of Learning.Checked.TargetGather: one thread per row,
row = block x block width + thread; a thread past the last row exits; the
others load the row's target t, the logit at row x vocabulary + t, and
store it to the row's slot. No arithmetic touches the word.
Preconditions (Realization contract):
the launch covers every row (blocks x block width >= rows);
the parameter block carries, at the three offsets given, the 64-bit
addresses of the output (4 bytes per row), the logits (row-major, 4
bytes each) and the targets (a 32-bit token per row);
rows x vocabulary x 4 fits 32 bits (the index is a 32-bit product);
every target is below the vocabulary.
The same generator makes Coppelius's image (output 0x160, logits 0x168,
targets 0x170, vocabulary 12288, 1024 rows in 4 blocks of 256) and the
checked path's (the linear step's layout: Checked.TargetGatherProbe).
28def targetGatherSM86CoalescedBlockThreads : Nat = 256The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.