
github.com
April 20, 2026
5 min read
54/100
Summary
A transformer model with approximately 25,000 parameters is implemented on an unmodified Commodore 64, utilizing hand-written 6502/6510 assembly. This 2-layer decoder-only architecture features real multi-head causal self-attention, softmax, and RMSNorm, and can be loaded from a floppy disk.
Key Takeaways
Community Sentiment
Positives
Concerns
Pure C, CPU-only inference with Mistral Voxtral Realtime 4B speech to text model
Feb 10, 2026

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Jul 31, 2026

An SLM trained on $8 ESP32-S3
Aug 5, 2026

Microgpt explained interactively
Mar 1, 2026

Nvidia PersonaPlex 7B on Apple Silicon: Full-Duplex Speech-to-Speech in Swift
Mar 5, 2026