9 articles · page 1 of 1
Kimi K2.7-Code: open-source coding model with better token efficiency
2026-06-12
Transformers Are Inherently Succinct
2026-06-05
Do transformers need three projections? Systematic study of QKV variants
2026-06-04
Soul Player C64 – A real transformer running on a 1 MHz Commodore 64
2026-04-20
Attention Residuals
2026-03-20
Building a Minimal Transformer for 10-digit Addition
2026-02-28
FORTH? Really!?
2026-02-06
Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation
2026-02-04
David Patterson: Challenges and Research Directions for LLM Inference Hardware
2026-01-25