The registers each instruction reads and writes, and the register demand
of a program: one owner for both the scoreboard pass (SM86.Scoreboard) and
the register count a launch declares. A 64-bit address or value occupies
the named register and the next; the tensor core's A, C and D fragments
are four registers from the named one and its B fragment two; LDSM writes
one, two or four. RZ (255) is not a register: it reads zero, absorbs
writes, and a fragment based at it names nothing.
The demand: the highest register named, plus one, plus TWO the hardware
reserves, rounded up to the allocation granule of eight. Found on an RTX
3070 (2026-09-23): the generated linear step at 2 x 6, 6 x 2 and 4 x 4 --
every shape whose declared count left exactly one register above the
highest named -- hung on the card with the GPU at 100 %, deterministically,
while every shape with two or more to spare ran; declaring the two
reserved registers made all three run. The machine model cannot see it.
---- one elimination per instruction ----
The schedule tools (SM86.Scoreboard's two passes,
Realization.Nvidia.SM86.StallCompaction) used to eliminate every
instruction body five or more times -- reads, writes, predicate writes,
latency class, fixed latency -- one 37-branch case split each. The
summary answers them all in one split: each field below is the same
branch body the corresponding function above uses, tupled, so every
field is definitionally the old answer and a pass walks each
instruction's body once instead of five times.
36constructor SM86OpSummaryValueThe compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.