github.com
March 22, 2026
6 min read
65/100
Summary
Flash-Moe is a pure C/Metal inference engine that runs the Qwen3.5-397B-A17B model, a 397 billion parameter Mixture-of-Experts model, on a MacBook Pro with 48GB RAM at over 4.4 tokens per second. The 209GB model streams from SSD using a custom Metal compute pipeline without relying on Python or other frameworks.
Key Takeaways
Community Sentiment
Positives
Concerns

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Jul 31, 2026

A 10 year old Xeon is all you need
Jun 1, 2026

DeepSeek V4 Flash on a Single AMD MI300X
Aug 4, 2026

TurboQuant KV Compression and SSD Expert Streaming for M5 Pro and IOS
Apr 1, 2026

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU
Jul 15, 2026