Source/Packages

Realization.Nvidia.SM86.StallCompaction

packages/realizations/cooperative/nvidia-sm86/src/Realization/Nvidia/SM86/StallCompaction.alpha

269 lines32 declarations15.0 KiBSHA-256 46e30c2c60d2

field · lines 41–41

compactionAnnotatedTail

Full file
A PROPOSER, not part of any argument: it rewrites a program's stall counts to the smallest each instruction's successor allows, by the same accounting SM86.Scoreboard's fixed-latency pass checks (issue one instruction, wait its stall; a fixed-latency result is ready sm86FixedLatency cycles after issue, a barrier sm86BarrierSetCycles after the instruction that sets it). Whatever it proposes is accepted only if the checker accepts it (SM86.LinearStepCheck.linearStepProgramAcceptedWith for the linear step); a mistake here is a rejected proposal, never a wrong program. Generated programs stall 15 cycles after every instruction (the maximum, chosen when nothing modelled latency); most instructions need 1, none less than its sm86MinimumStall. The annotated program: each instruction's schedule keys computed once. The old walk eliminated every body five times per step -- writes, latency and set keys for the issue at this step, reads, writes, guard, predicate and wait keys for the need of the previous one -- and rebuilt each instruction with a sixth split. Annotation does one 37-branch split per instruction up front (each field the same branch body the old helpers used), and carries the keys forward with the instruction itself, so the walk never re-examines a body except to rebuild it with its new stall.
41recursive unrestricted compactionAnnotatedTail

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.