The sm_121 realization of an sm_86 program: what a Blackwell GB10 runs for
it. Each instruction keeps its SM86 word (Accelerator.SM86's encoder)
wherever the two architectures mean the same thing; four rules rewrite the
rest, and a scheduler repairs the program's timing for sm_121's fixed
latencies:
* sm_121 has no constant-bank ALU operand: MOV Rd, c[b][o] becomes
LDC Rd, c[b][o], and IMAD / IMAD.WIDE with a constant operand become an
LDC (LDC.64) into a register followed by the register form. A constant
whose destination is also a source is loaded into a register (pair)
the program never names below its declared count, and the lowering
refuses when there is none. The LDC is variable-latency and signals
the program's free scoreboard.
* sm_86's RED is sm_121's REDG, with the descriptor and modifier bits
ptxas sets.
* Shared memory is based at 0x400 (the first kilobyte is reserved), so
every LDS, STS and LDSM offset moves up by it.
* The global-memory descriptor UR4:UR5 is zeroed in a two-instruction
prologue before the first access.
The scheduler walks the lowered instructions in order with the cycle each
is issued at. A register written by a fixed-latency instruction may be
read (or written again) only after the read-after-write latency of the
writer's class to the reader's (NVIDIA's sm100 table plus Mesa NAK's
sm_120 padding); a guard predicate after the predicate latency. The
instruction before is made to stall longer, up to fifteen cycles, and NOPs
carry the rest. Variable-latency results keep the program's own
scoreboards, except the LDCs above and a variable-latency instruction's
source registers that the SM86 program does not guard with a read
barrier: both go on the free scoreboard, and the next instruction touching
a loaded register or overwriting such a source waits on it. (Unguarded,
the block reductions' identity store -- STS of -inf, then the register
reused for the warp index -- stored 0 on sm_121 now and then: a causal
softmax row whose one score was far below zero summed to zero and came out
NaN.) A program with a branch is refused: its schedule would need the
latencies across the branch.
50constructor SM121LowerWideThe compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.