Source/Packages

Accelerator.SM86.Operands

packages/hardware/architectures/nvidia-sm86/src/Accelerator/SM86/Operands.alpha

645 lines51 declarations53.6 KiBSHA-256 e735be07682e

field · lines 43–43

sm86OpSummaryLatency

Full file
The registers each instruction reads and writes, and the register demand of a program: one owner for both the scoreboard pass (SM86.Scoreboard) and the register count a launch declares. A 64-bit address or value occupies the named register and the next; the tensor core's A, C and D fragments are four registers from the named one and its B fragment two; LDSM writes one, two or four. RZ (255) is not a register: it reads zero, absorbs writes, and a fragment based at it names nothing. The demand: the highest register named, plus one, plus TWO the hardware reserves, rounded up to the allocation granule of eight. Found on an RTX 3070 (2026-09-23): the generated linear step at 2 x 6, 6 x 2 and 4 x 4 -- every shape whose declared count left exactly one register above the highest named -- hung on the card with the GPU at 100 %, deterministically, while every shape with two or more to spare ran; declaring the two reserved registers made all three run. The machine model cannot see it. ---- one elimination per instruction ---- The schedule tools (SM86.Scoreboard's two passes, Realization.Nvidia.SM86.StallCompaction) used to eliminate every instruction body five or more times -- reads, writes, predicate writes, latency class, fixed latency -- one 37-branch case split each. The summary answers them all in one split: each field below is the same branch body the corresponding function above uses, tupled, so every field is definitionally the old answer and a pass walks each instruction's body once instead of five times.
43field unrestricted sm86OpSummaryLatency : Nat

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.