Source/Packages

Accelerator.SM121.Lowering

packages/hardware/architectures/nvidia-sm121/src/Accelerator/SM121/Lowering.alpha

2,621 lines365 declarations134.1 KiBSHA-256 b7b3bbc05e9c

def · lines 878–884

sm121LowerPlaceOf

Full file
2^bit for a bit below 64 without the doubling loop: 2^(bit mod 8) times 256^(bit div 8), each small power one byte shift and 256^q = ((2^q)^2)^2)^2. The loop costs a step per bit (~3.5k dispatch steps for high registers); mask insert and test call this per register, so the loop would dominate the once-per-program context folds that build the free scoreboard and scratch registers. (A chain of eight comparisons against bit div 8 took 93 steps a call, 8% of realizing a tiled product.)
878def sm121LowerPlaceOf =
879  (lambda unrestricted bit : Nat .
880    (let unrestricted low = (byte-to-nat (byte-shift-left (byte 1) (nat-to-byte (nat-modulo bit 8)))) in
881    (let unrestricted q = (byte-to-nat (byte-shift-left (byte 1) (nat-to-byte (nat-divide bit 8)))) in
882    (let unrestricted q2 = (nat-multiply q q) in
883    (let unrestricted q4 = (nat-multiply q2 q2) in
884      (nat-multiply low (nat-multiply q4 q4)))))))

The compiler supplied declaration spans and resolved links from this source snapshot. This page does not assert that this file belongs to a checked closure.