Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#ai-ethics#claude#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

Β© 2026 Themata.AI β€’ All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
πŸ•’ LatestπŸ”₯ Top
WeekMonthYearAll Time

Filtering by tag:

autonomous-modelsClear
NanoGPT Speedrun Frontier
nanogptllmsai-optimizationautonomous-models
Research

NanoGPT Speedrun Frontier

Prime Intellect ran 153 autonomous nanoGPT optimizer speedruns across 18 frontier-model entries and reported the best validated result for each. Fable 5, using Claude Code at the high setting, ranked first with a validated record of 2,726 and an 81.7% gap closed toward the stated human record of 2,600 from a baseline of 3,290. Its run used 3,010 steps and lasted 8.7 days. Opus 5 placed second at 2,920, closing 53.6% of the gap, followed by Kimi K3 at 2,930 with Prime Agent and 2,974 with Kimi Code. Opus 4.8 recorded 3,018, while GPT-5.6 Sol and GPT-5.6 Sol Pro reached 3,042 and 3,058. Sonnet 5 reached 3,105, and GPT-5.6 Luna reached 3,110. The remaining validated records ranged from 3,120 for Grok 4.5 and Qwen3.8 Max to 3,240 for Kimi K2.7. GLM 5.3 had no recorded result. The results use different agent harnesses and configurations, including Claude Code, Codex, Prime Agent, Kimi Code, Grok CLI, Qwen Code, Pi, and Muse Code. Prime Intellect also provides 41 curated full agent trajectories containing tool calls, subagents, and scratchpads.

primeintellect.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

4 min

12h ago

NanoGPT Speedrun Frontier

Prime Intellect ran 153 autonomous nanoGPT optimizer speedruns across 18 frontier-model entries and reported the best validated result for each. Fable 5, using Claude Code at the high setting, ranked first with a validated record of 2,726 and an 81.7% gap closed toward the stated human record of 2,600 from a baseline of 3,290. Its run used 3,010 steps and lasted 8.7 days. Opus 5 placed second at 2,920, closing 53.6% of the gap, followed by Kimi K3 at 2,930 with Prime Agent and 2,974 with Kimi Code. Opus 4.8 recorded 3,018, while GPT-5.6 Sol and GPT-5.6 Sol Pro reached 3,042 and 3,058. Sonnet 5 reached 3,105, and GPT-5.6 Luna reached 3,110. The remaining validated records ranged from 3,120 for Grok 4.5 and Qwen3.8 Max to 3,240 for Kimi K2.7. GLM 5.3 had no recorded result. The results use different agent harnesses and configurations, including Claude Code, Codex, Prime Agent, Kimi Code, Grok CLI, Qwen Code, Pi, and Muse Code. Prime Intellect also provides 41 curated full agent trajectories containing tool calls, subagents, and scratchpads.

primeintellect.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

4 min

12h ago

NanoGPT Speedrun Frontier

Prime Intellect ran 153 autonomous nanoGPT optimizer speedruns across 18 frontier-model entries and reported the best validated result for each. Fable 5, using Claude Code at the high setting, ranked first with a validated record of 2,726 and an 81.7% gap closed toward the stated human record of 2,600 from a baseline of 3,290. Its run used 3,010 steps and lasted 8.7 days. Opus 5 placed second at 2,920, closing 53.6% of the gap, followed by Kimi K3 at 2,930 with Prime Agent and 2,974 with Kimi Code. Opus 4.8 recorded 3,018, while GPT-5.6 Sol and GPT-5.6 Sol Pro reached 3,042 and 3,058. Sonnet 5 reached 3,105, and GPT-5.6 Luna reached 3,110. The remaining validated records ranged from 3,120 for Grok 4.5 and Qwen3.8 Max to 3,240 for Kimi K2.7. GLM 5.3 had no recorded result. The results use different agent harnesses and configurations, including Claude Code, Codex, Prime Agent, Kimi Code, Grok CLI, Qwen Code, Pi, and Muse Code. Prime Intellect also provides 41 curated full agent trajectories containing tool calls, subagents, and scratchpads.

primeintellect.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

4 min

12h ago

No more articles to load