The scoreboard pass: the machine-level check the functional model cannot
make. SM86.MachineModel gives every instruction its result at once; the
silicon does not. A variable-latency instruction (S2R, a global load,
shuffles, shared-memory and tensor-core forms) delivers its destination
later, and the only thing that orders a consumer after it is the control
word: the producer names a write barrier (SB0..SB5) and the consumer's
wait mask names that barrier. A consumer without the wait reads whatever
the register held before -- on the RTX 3090 (2026-09-22) the first linear
step realization read zeros for W, x and t and wrote a record of zeros,
while the functional model had accepted it.
This pass walks the program in order as one straight-line issue stream
(predication does not change issue), keeping the registers whose value is
still in flight and the barrier each waits on. An instruction first
retires every pending register whose barrier is in its wait mask, then any
register it reads or writes that is still pending is a hazard, then its
own variable-latency destinations become pending on its write barrier -- or
on a barrier no mask can name, when it declares none, so every later use
is a hazard. The result is 0 for a clean program, else the ordinal (from
1) of the first hazardous instruction. Fixed-latency forms are ordered
by their stall counts, which this pass does not model.
37field unrestricted scoreboardPendingBarrier : NatThe compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.