
primeintellect.ai
August 22, 2026
4 min read
48/100
Summary
Prime Intellect ran 153 autonomous nanoGPT optimizer speedruns across 18 frontier-model entries and reported the best validated result for each. Fable 5, using Claude Code at the high setting, ranked first with a validated record of 2,726 and an 81.7% gap closed toward the stated human record of 2,600 from a baseline of 3,290. Its run used 3,010 steps and lasted 8.7 days. Opus 5 placed second at 2,920, closing 53.6% of the gap, followed by Kimi K3 at 2,930 with Prime Agent and 2,974 with Kimi Code. Opus 4.8 recorded 3,018, while GPT-5.6 Sol and GPT-5.6 Sol Pro reached 3,042 and 3,058. Sonnet 5 reached 3,105, and GPT-5.6 Luna reached 3,110. The remaining validated records ranged from 3,120 for Grok 4.5 and Qwen3.8 Max to 3,240 for Kimi K2.7. GLM 5.3 had no recorded result. The results use different agent harnesses and configurations, including Claude Code, Codex, Prime Agent, Kimi Code, Grok CLI, Qwen Code, Pi, and Muse Code. Prime Intellect also provides 41 curated full agent trajectories containing tool calls, subagents, and scratchpads.
Key Takeaways
What the discussion said
The thread spent less time celebrating a frontier-model leaderboard than trying to decode what the benchmark actually measures. Several readers found the write-up needlessly opaque: it did not plainly define an autonomous run, the optimizer objective, frozen verification, or why repeated fixed-seed trials represent research ability. Others reconstructed the setup as agents iteratively altering a small model’s training recipe, testing it against a target loss, and retaining successful changes. That made the evaluation sound genuinely useful rather than mystical, but commenters felt the article should have done that explanatory work itself. There is real interest in the task as an AI R&D proxy. Readers liked that it turns optimization work into a concrete, automatically checkable competition, and one low-cost model’s near-Sonnet result was taken as evidence that cheap, fast agents can be highly effective when verification is crisp. The poor result from one competing model drew attention, though the discussion resisted treating it as a clean verdict on raw intelligence. Long-horizon agent training, harness quality, waiting behavior, effort settings, prompt wording, and nonidentical experimental configurations may all move the score. Commenters also wondered whether stronger memory for tentative findings would separate agents more than their ability to discover the same obvious ideas. Overall, the benchmark is seen as promising, but its presentation and comparability need tightening before model rankings deserve much confidence.
Where opinion split
The main dispute is whether the leaderboard reflects models' autonomous research ability or mostly the quality and configuration of the agent harness. Supporters see a difficult, repeatable optimize-test loop with automatic validation as a meaningful R&D task; skeptics argue that unequal settings, long-horizon training exposure, and weak experimental memory can dominate the apparent model differences.
Community Sentiment
Positives
Concerns