
w4g1.dev
August 16, 2026
6 min read
63/100
Summary
Reasoning scores for AI models are increasing while per-token compute is decreasing. GLM-5.2 achieves 99.2% on AIME 2026 with 40 billion parameters, Qwen3.5 scores 91.3% with 17 billion parameters, and DeepSeek V4-Flash operates with 13 billion parameters, contrasting with GPT-4's rumored 280 billion parameters which struggled with AIME problems.
Key Takeaways
Community Sentiment
Positives
Concerns

Qwen 3.6 27B is the sweet spot for local development
Jun 29, 2026

Local Qwen isn't a worse Opus, it's a different tool
Jun 18, 2026

The session you cannot take with you
Jul 31, 2026

Qwen 3.8 27B is excellent, but it defaults to overthinking things
Aug 16, 2026

What Happens When the Cost of Intelligence Drops 100x
Aug 21, 2026