
github.com
September 8, 2026
10 min read
60/100
Summary
TL;DR — Kimi K3 (2.8T-parameter MoE, 1.45 TB of expert weights) running on one M5 Max MacBook Pro with 128 GB, experts streamed from four SSDs. - 1.00 tok/s steady decode over a 512-token answer; 1.13 over 128; 0.96 on the public 17-token prompt (upstream reported 0.68). - The honest limit: a 512-token prompt takes ~6.3 minutes to its first token. Cause found (prefill re-reads each layer's experts...