
catalystneuro.com
August 21, 2026
17 min read
51/100
Summary
CatalystNeuro founder Ben Dichter analyzed Artificial Analysis benchmark data and found that the cheapest measured cost for a given level of large-language-model capability has fallen sharply. Models scoring at least 40 on the Artificial Analysis Intelligence Index fell from $1.22 per evaluated task in February 2026 to $0.022 as of August 19, a 56-fold decline. The index combines nine evaluations weighted toward agentic tasks, coding, scientific reasoning, and general capability. GPT-5.6 Luna’s effort settings covered much of the lower-cost frontier, while Claude Opus 5 reached the highest cited score, 63.1, at $2.34 per task at maximum effort. The analysis estimates that cost records for capability tiers of 40, 50, and 60 or higher have been halving roughly every four to ten weeks, though historical price cuts and incomplete retired-model data limit the estimates. Dichter predicts that, if the trend holds, models at index 60 could cost under $0.10 per task within a couple of quarters. Lower costs can make large-scale work such as literature reviews, legal discovery, data curation, moderation, and support triage economically viable. Dichter argues that cheaper model calls may increase total AI spending because organizations can run full-corpus, repeated, and consensus-based workflows that were previously too expensive. OpenRouter offers routing based on a minimum capability score to select the cheapest qualifying frontier model.
Key Takeaways
What the discussion said
Commenters largely accepted the article’s central observation: the price of a given level of model capability is collapsing fast, and the custom historical price-versus-capability charts made that trend unusually hard to dismiss. Several readers argued that this shift matters less as a cheaper chatbot bill than as an unlock for workloads previously too expensive to attempt: exhaustive document review, cheap domain models, and robotics whose perception-and-planning loops currently move at a painfully slow pace. But the thread pushed back on treating token price as the true price of intelligence. Lower per-token costs may invite vastly more inference, longer chains of reasoning, consensus voting, and agent loops, so total spend need not fall. Others stressed that latency and reliability are binding constraints: a low-cost model that pauses for half a minute or cannot sustain a real codebase is not interchangeable with a pricier frontier system. Small Chinese and open models earned genuine praise for delivering useful work at a fraction of frontier pricing, though readers expect frontier labs to keep improving too. The practical consensus was that cost curves are extraordinary but incomplete. Hardware, distillation, and competition could drive another major drop, perhaps eventually enabling strong local models, while sensor quality, memory, battery limits, training subsidies, and uncertain demand elasticity make confident economic forecasts premature.
Where opinion split
The sharpest dispute was whether plunging token prices genuinely make intelligence cheap. Optimists argued that capable open and small models already perform valuable tasks for tiny sums, and falling costs will unlock nearly unlimited new software and automation demand. Skeptics replied that usage will expand to consume the savings, while slow, unreliable outputs and subsidized inference mean advertised token prices are not yet a clean measure of usable intelligence.
Community Sentiment
Positives
Concerns