
blog.cloudflare.com
August 3, 2026
7 min read
53/100
Summary
Workers AI runs inference for Moonshot's Kimi K-series and Z.ai's GLM models on GPUs in Cloudflare data centers. These models are large, long-context, and mixture-of-experts types, presenting challenges in efficient serving due to memory constraints.
Key Takeaways
Community Sentiment
Positives
Concerns

Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
May 29, 2026

A 10 year old Xeon is all you need
Jun 1, 2026

Running a One Trillion-Parameter LLM Locally on AMD Ryzen AI Max+ Cluster
Mar 1, 2026

Performance per dollar is getting faster and cheaper
Jul 3, 2026

Local Qwen isn't a worse Opus, it's a different tool
Jun 18, 2026