Coppelius
A 57.7M-parameter transformer: 16 blocks, width 512, rotary attention, gated GELU, bfloat16 matrix products and AdamW, trained on WikiText-103.
- Runs on
- x86-64 or AArch64 host → sm_86 (RTX 3070, 3090) and sm_121 (GB10); sm_89 envelope for the RTX 4090
- Status
- Trains, resumes from a checkpoint and predicts on the RTX 3070, RTX 3090 and DGX Spark GB10. On the RTX 3090 it ran 10,000 steps and trained 1.51× faster than PyTorch with CUDA graphs in matched runs. A fresh RTX 4090 run is pending.
- Namespaces
- Coppelius · Training