Schema enables frontier models to function like physicists by allowing them to write executable programs for game mechanisms, test these programs against reality, and plan within them. This method has achieved approximately 99% accuracy on the ARC-AGI-3 public benchmark.
schema-harness.github.io
17 min
7/16/2026
The Agentica SDK by Symbolica achieves a score of 36.08% on ARC-AGI-3, completing 113 out of 182 levels and 7 out of 25 games. This performance surpasses CoT baselines significantly, with Agentica costing $1,005 compared to Opus 4.6's $8,900 for a mere 0.25% score.
symbolica.ai
1 min
3/27/2026
Schema enables frontier models to function like physicists by allowing them to write executable programs for game mechanisms, test these programs against reality, and plan within them. This method has achieved approximately 99% accuracy on the ARC-AGI-3 public benchmark.
schema-harness.github.io
17 min
7/16/2026
The Agentica SDK by Symbolica achieves a score of 36.08% on ARC-AGI-3, completing 113 out of 182 levels and 7 out of 25 games. This performance surpasses CoT baselines significantly, with Agentica costing $1,005 compared to Opus 4.6's $8,900 for a mere 0.25% score.
symbolica.ai
1 min
3/27/2026
Schema enables frontier models to function like physicists by allowing them to write executable programs for game mechanisms, test these programs against reality, and plan within them. This method has achieved approximately 99% accuracy on the ARC-AGI-3 public benchmark.
schema-harness.github.io
17 min
7/16/2026
The Agentica SDK by Symbolica achieves a score of 36.08% on ARC-AGI-3, completing 113 out of 182 levels and 7 out of 25 games. This performance surpasses CoT baselines significantly, with Agentica costing $1,005 compared to Opus 4.6's $8,900 for a mere 0.25% score.
symbolica.ai
1 min
3/27/2026
No more articles to load