Source/Packages

Accelerator.SM86.Operands

packages/hardware/architectures/nvidia-sm86/src/Accelerator/SM86/Operands.alpha

645 lines51 declarations53.6 KiBSHA-256 e735be07682e

constructor · lines 36–36

SM86OpSummaryValue

Full file
The registers each instruction reads and writes, and the register demand of a program: one owner for both the scoreboard pass (SM86.Scoreboard) and the register count a launch declares. A 64-bit address or value occupies the named register and the next; the tensor core's A, C and D fragments are four registers from the named one and its B fragment two; LDSM writes one, two or four. RZ (255) is not a register: it reads zero, absorbs writes, and a fragment based at it names nothing. The demand: the highest register named, plus one, plus TWO the hardware reserves, rounded up to the allocation granule of eight. Found on an RTX 3070 (2026-09-23): the generated linear step at 2 x 6, 6 x 2 and 4 x 4 -- every shape whose declared count left exactly one register above the highest named -- hung on the card with the GPU at 100 %, deterministically, while every shape with two or more to spare ran; declaring the two reserved registers made all three run. The machine model cannot see it. ---- one elimination per instruction ---- The schedule tools (SM86.Scoreboard's two passes, Realization.Nvidia.SM86.StallCompaction) used to eliminate every instruction body five or more times -- reads, writes, predicate writes, latency class, fixed latency -- one 37-branch case split each. The summary answers them all in one split: each field below is the same branch body the corresponding function above uses, tupled, so every field is definitionally the old answer and a pass walks each instruction's body once instead of five times.
36constructor SM86OpSummaryValue

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.