The future of business, today.
RSSNewslettersAdvertise
Business Future Today

Model economics

GPT-6 Astra’s Coding Gains Put Token Efficiency at the Center of Model Buying

Artificial Analysis finds GPT-6 Astra matching leading coding-agent scores at materially lower task cost—but its higher list price weakens the case for general intelligence workloads.

Editorial image for GPT-6 Astra’s Coding Gains Put Token Efficiency at the Center of Model Buying
Hacker News

GPT-6 Astra has made a meaningful move in coding-agent economics, according to new [Artificial Analysis benchmarking](https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra). The model scored 67 in the firm’s Coding Agent Index, roughly level with Claude Opus 5, Fable 5 and Muse Spark 1.3 in their respective coding harnesses. Fable 5.1 remains the reported leader at 70.

The more consequential result for engineering leaders is not the headline score. It is the amount of model output needed to reach it.

A better coding cost curve

Artificial Analysis says GPT-6 Astra, at its maximum effort setting, uses about one-third of the tokens consumed by GPT-5.6 Sol at maximum effort in its Codex harness. It also reports that Astra uses about one-fifth as many tokens as Claude Opus 5 at its xhigh setting.

That reduction changes the economics despite a sharp price increase. Astra is priced at $10 per million input tokens and $50 per million output tokens, versus $4 and $20 for GPT-5.6 Sol—2.5 times higher across both categories. Cache-read and cache-write pricing adjustments remain a 90% discount and 25% premium, respectively.

Supporting image for GPT-6 Astra’s Coding Gains Put Token Efficiency at the Center of Model Buying
Hacker News

For coding tasks, the token savings appear large enough to offset the new rates. Artificial Analysis places Astra on the cost-efficiency frontier of its Coding Agent Index: at max effort, it costs about the same per task as GPT-5.6 Sol while scoring two points higher. It estimates Astra costs less than half as much per task as Fable 5 for the same index score.

For teams running repository-scale refactors, test generation, bug investigation or autonomous code workflows, this is the practical signal: list token prices are no longer sufficient for vendor selection. The token footprint of completed work—and the retry rate needed to get reliable patches—can dominate the bill.

The gains do not extend cleanly to general work

Astra’s position is less favorable in Artificial Analysis’s broader Intelligence Index. It scored 61, equal to GPT-5.6 Sol and five points behind Claude Fable 5.1 at maximum effort with fallback. It also trailed Muse Spark 1.3.

The model used roughly 10% fewer output tokens than its predecessor at max effort, but that was not enough to neutralize the 2.5x price rise. Artificial Analysis estimates Astra is 75% more expensive per task than GPT-5.6 Sol in this setting and sits largely behind its predecessor on the intelligence-cost frontier.

That distinction matters for companies standardizing on one provider or one default model. Astra may be a compelling routing target for coding agents while being a weaker default for general research, analysis and knowledge-work automation on a cost-adjusted basis.

Reliability improved, but evaluation results remain mixed

The benchmarks also point to a notable reliability improvement. On Artificial Analysis’s AA-Omniscience knowledge and hallucination evaluation, Astra’s reported hallucination rate fell from 92% to 51% at max effort, while accuracy rose by four points.

It gained roughly 80 Elo points on AA-Briefcase, a long-horizon knowledge-work evaluation involving linked tasks and large source collections. But the report also found an approximately 80-point decline on GDPval-AA v2, which measures economically valuable tasks across 44 occupations. It cited smaller regressions in customer support, scientific Python and long-context reasoning tests, as well as lower presentation-quality scores in AA-Briefcase.

What to watch next

Operators should test Astra on their own task mix rather than extrapolating from a single aggregate ranking. Track completed-task cost, latency, tool-call and retry behavior, patch acceptance, and human review burden. The key question is not whether Astra is broadly “better,” but whether its coding token efficiency survives in a production agent loop.

The larger market lesson is that benchmark leadership is fragmenting by workload. Model buyers will increasingly need routing policies that optimize separately for coding, long-horizon research, customer operations and document-heavy reasoning—rather than treating one model score as a universal purchasing signal.

Sources

STAY AHEAD

The future of business, in your inbox.

Useful signals on the companies, technologies and shifts changing business.

One useful briefing. Unsubscribe any time.