Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#discussion#anthropic

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
trendingdiscussion

Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs

GitHub - argonautlabsai/deltafin: ARGODRIVE Deltafin: Kimi K3 (2.8T MoE) streamed from SSDs on Apple Silicon — fork of gavamedia/deltafin with the ARGODRIVE storage work and benchmark package

github.com

September 8, 2026

10 min read

🔥🔥🔥🔥🔥

60/100

Summary

TL;DR — Kimi K3 (2.8T-parameter MoE, 1.45 TB of expert weights) running on one M5 Max MacBook Pro with 128 GB, experts streamed from four SSDs. - 1.00 tok/s steady decode over a 512-token answer; 1.13 over 128; 0.96 on the public 17-token prompt (upstream reported 0.68). - The honest limit: a 512-token prompt takes ~6.3 minutes to its first token. Cause found (prefill re-reads each layer's experts...

Read original article