Training
Models training on real hardware.
Each dashboard reads the run's own records: loss at every step, predictions, where each step's device time goes, optimizer internals, and GPU and host telemetry. The executables come from the Alpha compiler alone, with no CUDA runtime in the path.
Coppelius on the RTX 3090
a rented RTX 3090 (RunPod) · sm_86
2,000 of 2,000 steps · updated 2026-09-27 22:30 UTC
Open dashboard →FinishedCoppelius on the GB10
the DGX Spark · sm_121
1,000 of 1,000 invocations · updated 2026-09-27 01:43 UTC
Open dashboard →Not startedBaguette on the RTX 3090
a rented RTX 3090 (RunPod) · sm_86
The dashboard is ready; no run has published a snapshot yet.
Open dashboard →Milestones
Recorded results
Each links to the run record it summarizes.
- Alpha trains Coppelius 1.51× faster than PyTorch with CUDA graphs
Matched 2,000-step runs from the same checkpoint on an RTX 3090: 35,683 tokens/s against 23,675, with mean losses within 0.004 nats.
- 10,000 steps on the RTX 3090
One process, ten checkpoints, every loss finite, then a fresh-process resume. 36,310 tokens/s wall.
- A 45 ms training step on the DGX Spark
Fused kernels and a single optimizer launch bring the GB10 step to 45.35 ms, with losses and checkpoints byte-identical to the previous run.
- Coppelius trains on Blackwell
The compiler lowers the sm_86 kernels to sm_121 and Coppelius trains on the GB10 for 10,000 steps, then resumes and predicts.