Source/Packages

Accelerator.SM121.Lowering

packages/hardware/architectures/nvidia-sm121/src/Accelerator/SM121/Lowering.alpha

2,621 lines365 declarations134.1 KiBSHA-256 b7b3bbc05e9c

constructor · lines 53–53

SM121LowerVariable

Full file
The sm_121 realization of an sm_86 program: what a Blackwell GB10 runs for it. Each instruction keeps its SM86 word (Accelerator.SM86's encoder) wherever the two architectures mean the same thing; four rules rewrite the rest, and a scheduler repairs the program's timing for sm_121's fixed latencies: * sm_121 has no constant-bank ALU operand: MOV Rd, c[b][o] becomes LDC Rd, c[b][o], and IMAD / IMAD.WIDE with a constant operand become an LDC (LDC.64) into a register followed by the register form. A constant whose destination is also a source is loaded into a register (pair) the program never names below its declared count, and the lowering refuses when there is none. The LDC is variable-latency and signals the program's free scoreboard. * sm_86's RED is sm_121's REDG, with the descriptor and modifier bits ptxas sets. * Shared memory is based at 0x400 (the first kilobyte is reserved), so every LDS, STS and LDSM offset moves up by it. * The global-memory descriptor UR4:UR5 is zeroed in a two-instruction prologue before the first access. The scheduler walks the lowered instructions in order with the cycle each is issued at. A register written by a fixed-latency instruction may be read (or written again) only after the read-after-write latency of the writer's class to the reader's (NVIDIA's sm100 table plus Mesa NAK's sm_120 padding); a guard predicate after the predicate latency. The instruction before is made to stall longer, up to fifteen cycles, and NOPs carry the rest. Variable-latency results keep the program's own scoreboards, except the LDCs above and a variable-latency instruction's source registers that the SM86 program does not guard with a read barrier: both go on the free scoreboard, and the next instruction touching a loaded register or overwriting such a source waits on it. (Unguarded, the block reductions' identity store -- STS of -inf, then the register reused for the warp index -- stored 0 on sm_121 now and then: a causal softmax row whose one score was far below zero summed to zero and came out NaN.) A program with a branch is refused: its schedule would need the latencies across the branch.
53constructor SM121LowerVariable

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.