The registers each instruction reads and writes, and the register demand
of a program: one owner for both the scoreboard pass (SM86.Scoreboard) and
the register count a launch declares. A 64-bit address or value occupies
the named register and the next; the tensor core's A, C and D fragments
are four registers from the named one and its B fragment two; LDSM writes
one, two or four. RZ (255) is not a register: it reads zero, absorbs
writes, and a fragment based at it names nothing.
The demand: the highest register named, plus one, plus TWO the hardware
reserves, rounded up to the allocation granule of eight. Found on an RTX
3070 (2026-09-23): the generated linear step at 2 x 6, 6 x 2 and 4 x 4 --
every shape whose declared count left exactly one register above the
highest named -- hung on the card with the GPU at 100 %, deterministically,
while every shape with two or more to spare ran; declaring the two
reserved registers made all three run. The machine model cannot see it.
---- one elimination per instruction ----
The schedule tools (SM86.Scoreboard's two passes,
Realization.Nvidia.SM86.StallCompaction) used to eliminate every
instruction body five or more times -- reads, writes, predicate writes,
latency class, fixed latency -- one 37-branch case split each. The
summary answers them all in one split: each field below is the same
branch body the corresponding function above uses, tupled, so every
field is definitionally the old answer and a pass walks each
instruction's body once instead of five times.
47field unrestricted sm86OpSummaryControl : (family SM86Control)The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.