Prime Intellect ran 153 autonomous nanoGPT optimizer speedruns across 18 frontier-model entries and reported the best validated result for each. Fable 5, using Claude Code at the high setting, ranked first with a validated record of 2,726 and an 81.7% gap closed toward the stated human record of 2,600 from a baseline of 3,290. Its run used 3,010 steps and lasted 8.7 days. Opus 5 placed second at 2,920, closing 53.6% of the gap, followed by Kimi K3 at 2,930 with Prime Agent and 2,974 with Kimi Code. Opus 4.8 recorded 3,018, while GPT-5.6 Sol and GPT-5.6 Sol Pro reached 3,042 and 3,058. Sonnet 5 reached 3,105, and GPT-5.6 Luna reached 3,110. The remaining validated records ranged from 3,120 for Grok 4.5 and Qwen3.8 Max to 3,240 for Kimi K2.7. GLM 5.3 had no recorded result. The results use different agent harnesses and configurations, including Claude Code, Codex, Prime Agent, Kimi Code, Grok CLI, Qwen Code, Pi, and Muse Code. Prime Intellect also provides 41 curated full agent trajectories containing tool calls, subagents, and scratchpads.
primeintellect.ai
4 min
8/22/2026
A follow-up question on a 50,000 token conversation with an AI can take three to five minutes before the first token is generated. A three-week effort was dedicated to debugging a cache leak in order to optimize performance on the Mac Studio.
mrzk.io
12 min
7/11/2026
TensorZero is an open-source LLMOps platform that provides a unified API for accessing various LLM providers with less than 1ms p99 latency. It includes features for observability, evaluation, optimization, and experimentation, allowing users to store inferences, benchmark workflows, optimize prompts, and manage experiments programmatically.
github.com
7 min
6/13/2026
FlashAttention-T introduces a fully tensorized attention mechanism that leverages tensor-vector parallelism to enhance performance. This innovation aims to improve the efficiency of attention-based models in various applications.
dl.acm.org
6 min
2/3/2026
Prime Intellect ran 153 autonomous nanoGPT optimizer speedruns across 18 frontier-model entries and reported the best validated result for each. Fable 5, using Claude Code at the high setting, ranked first with a validated record of 2,726 and an 81.7% gap closed toward the stated human record of 2,600 from a baseline of 3,290. Its run used 3,010 steps and lasted 8.7 days. Opus 5 placed second at 2,920, closing 53.6% of the gap, followed by Kimi K3 at 2,930 with Prime Agent and 2,974 with Kimi Code. Opus 4.8 recorded 3,018, while GPT-5.6 Sol and GPT-5.6 Sol Pro reached 3,042 and 3,058. Sonnet 5 reached 3,105, and GPT-5.6 Luna reached 3,110. The remaining validated records ranged from 3,120 for Grok 4.5 and Qwen3.8 Max to 3,240 for Kimi K2.7. GLM 5.3 had no recorded result. The results use different agent harnesses and configurations, including Claude Code, Codex, Prime Agent, Kimi Code, Grok CLI, Qwen Code, Pi, and Muse Code. Prime Intellect also provides 41 curated full agent trajectories containing tool calls, subagents, and scratchpads.
primeintellect.ai
4 min
8/22/2026
Mag 7 is experiencing a decline in performance metrics. Data indicates a need for strategic adjustments to improve outcomes.
apollo.com
1 min
6/29/2026
A follow-up question on a 50,000 token conversation with an AI can take three to five minutes before the first token is generated. A three-week effort was dedicated to debugging a cache leak in order to optimize performance on the Mac Studio.
mrzk.io
12 min
7/11/2026
TensorZero is an open-source LLMOps platform that provides a unified API for accessing various LLM providers with less than 1ms p99 latency. It includes features for observability, evaluation, optimization, and experimentation, allowing users to store inferences, benchmark workflows, optimize prompts, and manage experiments programmatically.
github.com
7 min
6/13/2026
FlashAttention-T introduces a fully tensorized attention mechanism that leverages tensor-vector parallelism to enhance performance. This innovation aims to improve the efficiency of attention-based models in various applications.
dl.acm.org
6 min
2/3/2026
Prime Intellect ran 153 autonomous nanoGPT optimizer speedruns across 18 frontier-model entries and reported the best validated result for each. Fable 5, using Claude Code at the high setting, ranked first with a validated record of 2,726 and an 81.7% gap closed toward the stated human record of 2,600 from a baseline of 3,290. Its run used 3,010 steps and lasted 8.7 days. Opus 5 placed second at 2,920, closing 53.6% of the gap, followed by Kimi K3 at 2,930 with Prime Agent and 2,974 with Kimi Code. Opus 4.8 recorded 3,018, while GPT-5.6 Sol and GPT-5.6 Sol Pro reached 3,042 and 3,058. Sonnet 5 reached 3,105, and GPT-5.6 Luna reached 3,110. The remaining validated records ranged from 3,120 for Grok 4.5 and Qwen3.8 Max to 3,240 for Kimi K2.7. GLM 5.3 had no recorded result. The results use different agent harnesses and configurations, including Claude Code, Codex, Prime Agent, Kimi Code, Grok CLI, Qwen Code, Pi, and Muse Code. Prime Intellect also provides 41 curated full agent trajectories containing tool calls, subagents, and scratchpads.
primeintellect.ai
4 min
8/22/2026
TensorZero is an open-source LLMOps platform that provides a unified API for accessing various LLM providers with less than 1ms p99 latency. It includes features for observability, evaluation, optimization, and experimentation, allowing users to store inferences, benchmark workflows, optimize prompts, and manage experiments programmatically.
github.com
7 min
6/13/2026
A follow-up question on a 50,000 token conversation with an AI can take three to five minutes before the first token is generated. A three-week effort was dedicated to debugging a cache leak in order to optimize performance on the Mac Studio.
mrzk.io
12 min
7/11/2026
Mag 7 is experiencing a decline in performance metrics. Data indicates a need for strategic adjustments to improve outcomes.
apollo.com
1 min
6/29/2026
FlashAttention-T introduces a fully tensorized attention mechanism that leverages tensor-vector parallelism to enhance performance. This innovation aims to improve the efficiency of attention-based models in various applications.
dl.acm.org
6 min
2/3/2026
No more articles to load