OpenAI Ships GPT-6 Astra: 99.9% on ARC-AGI-3 and a Head-On Fight with Claude Fable
OpenAI has begun rolling out GPT-6 Astra, a model priced to directly undercut Claude Fable while claiming a near-perfect 99.9% score on the ARC-AGI-3 benchmark.
GPT-6 Astra started limited distribution to organizations on September 3, 2026, with broader availability for ChatGPT Plus, Pro, Business, and Enterprise users following in subsequent days. The model is accessible via the OpenAI API and AWS under the identifier gpt-6-astra. Pricing is set at $10 per million input tokens and $50 per million output tokens, matching the rate structure of Anthropic's Claude Fable 5 and 5.1. This parity positions Astra as a direct market competitor, though OpenAI claims superior performance on most self-reported benchmarks.
The headline metric is a 99.9% score on the ARC-AGI-3 benchmark, released in March 2026. However, this result required OpenAI's custom "Provider Adapter harness," which preserves opaque reasoning state between requests and uses compaction for longer conversations to reuse prior work; this run cost $19,000. By contrast, the default ARC-AGI harness yielded a 62.7% score at a cost of $26,000. No published result exists yet for Claude Fable 5 on this specific test. In security evaluations, Astra achieved 100% on ExploitBench, compared to 78.5% for GPT-5.6 Sol, and reached 99.2% success within four attempts on SRE-Bench binary reverse engineering, significantly outperforming Sol's 68.7%.
Performance diverges when examining general intelligence versus coding efficiency. Artificial Analysis reports that Astra scores 61 on their Intelligence Index, equal to GPT-5.6 Sol but five points behind Claude Fable 5.1 (max with fallback) and trailing Meta's Muse Spark 1.3. Conversely, Astra leads the Coding Agent Index cost efficiency frontier. At maximum effort, it costs approximately the same as GPT-5.6 Sol while scoring two points higher. Per task, the model operates at less than half the cost of Claude Fable 5 for an equivalent score. Long context capabilities also show improvement, with Astra achieving 100% accuracy on OpenAI's eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens.