The fewest cycles an instruction must stall before the next issues, from
its own issue rather than a result: 1, except BAR.SYNC.DEFER_BLOCKING
(6) and an HMMA ordered by its fixed latency (8). ptxas always gives
the barrier 6 (sm_86 and sm_121, every bar.sync read off nvdisasm -hex on
the DGX Spark, 2026-09-26). With 1 the GB10 let a warp's
shared-matrix load right after the barrier read, now and then, what
another warp had stored before it: the tiled product's last k-step (its
loads follow the barrier directly) read the buffer's previous tile.
487def sm86BarrierSynchronizeStall : Nat = 6The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.