Source/Packages

Accelerator.SM86.Operands

packages/hardware/architectures/nvidia-sm86/src/Accelerator/SM86/Operands.alpha

645 lines51 declarations53.6 KiBSHA-256 e735be07682e

def · lines 487–487

sm86BarrierSynchronizeStall

Full file
The fewest cycles an instruction must stall before the next issues, from its own issue rather than a result: 1, except BAR.SYNC.DEFER_BLOCKING (6) and an HMMA ordered by its fixed latency (8). ptxas always gives the barrier 6 (sm_86 and sm_121, every bar.sync read off nvdisasm -hex on the DGX Spark, 2026-09-26). With 1 the GB10 let a warp's shared-matrix load right after the barrier read, now and then, what another warp had stored before it: the tiled product's last k-step (its loads follow the barrier directly) read the buffer's previous tile.
487def sm86BarrierSynchronizeStall : Nat = 6

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.