DeepSeek V4 Flash 0731 achieves scores of 89.0% on ARC-AGI-1 Semi-Private at a cost of $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task. The model also records scores of 87.0% and 56.0% for the High variant, and 84.0% and 46.0% for the Low variant on the respective benchmarks.
arcprize.org
21 min
8/7/2026
ARC-AGI-3 measures AI agents' ability to adapt to novel interactive environments, evolving from previous versions that focused on passive fluid intelligence. The leaderboard visualizes the relationship between cost-per-task and performance, emphasizing efficiency in problem-solving.
arcprize.org
1 min
7/25/2026
DeepSeek V4 Flash 0731 achieves scores of 89.0% on ARC-AGI-1 Semi-Private at a cost of $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task. The model also records scores of 87.0% and 56.0% for the High variant, and 84.0% and 46.0% for the Low variant on the respective benchmarks.
arcprize.org
21 min
8/7/2026
ARC-AGI-3 measures AI agents' ability to adapt to novel interactive environments, evolving from previous versions that focused on passive fluid intelligence. The leaderboard visualizes the relationship between cost-per-task and performance, emphasizing efficiency in problem-solving.
arcprize.org
1 min
7/25/2026
DeepSeek V4 Flash 0731 achieves scores of 89.0% on ARC-AGI-1 Semi-Private at a cost of $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task. The model also records scores of 87.0% and 56.0% for the High variant, and 84.0% and 46.0% for the Low variant on the respective benchmarks.
arcprize.org
21 min
8/7/2026
ARC-AGI-3 measures AI agents' ability to adapt to novel interactive environments, evolving from previous versions that focused on passive fluid intelligence. The leaderboard visualizes the relationship between cost-per-task and performance, emphasizing efficiency in problem-solving.
arcprize.org
1 min
7/25/2026
No more articles to load