New Horizon No. 247 / 2026-09-04 · Berlin

OpenAI's new flagship model matches Claude Fable's pricing at $10/$50 per million tokens, though its headline benchmark result depends on a proprietary harness that scored 62.7% under the default ARC-AGI setup.
Generated via ComfyUI / Z-Image Turbo

What shipped

OpenAI shipped GPT-6 Astra on 3 September 2026, after a launch delay that released embargoed benchmark figures early, according to wccftech.com. The model is rolling out first to a limited set of organizations, then over the coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, with access through the OpenAI API and AWS. Fortune reports OpenAI is also promoting computer use — the model operating a user's machine — as a headline capability.

API pricing matches Claude Fable 5 and 5.1 exactly: $10 per million input tokens and $50 per million output tokens. Simon Willison, writing on simonwillison.net, describes Astra as OpenAI's direct Fable competitor and notes it appears to score higher than Fable on most of OpenAI's self-reported benchmarks. Fable 5 has no published ARC-AGI-3 result, so the comparison on that benchmark currently runs one way.

Matching Fable's rate card removes price as a differentiator between the two flagships. That suggests OpenAI intends to compete on measured capability alone, which raises the evidentiary weight of those self-reported numbers. Willison states he had not tried the model himself at the time of writing, so early coverage rests on OpenAI's own figures rather than independent hands-on testing by early users.

The harness question

The headline number is 99.9 percent on ARC-AGI-3, a benchmark released in March. The ARC-AGI blog, as cited by Willison, notes the score was achieved for $19,000 using OpenAI's custom Provider Adapter harness. Under the default ARC-AGI harness, the same model scored 62.7 percent at a cost of $26,000. Both figures come from the benchmark's maintainers rather than from OpenAI's marketing materials.

The Provider Adapter harness preserves opaque reasoning state, per the ARC-AGI blog; the published record cuts off before describing what it preserves that state between or how. What the two numbers establish is a 37.2-point spread between a vendor-built harness and the benchmark's default setup, with the custom harness also cheaper to run. The mechanism details matter because reasoning-state persistence is exactly the kind of scaffolding that can change task outcomes.

A 99.9 percent score on a vendor-built harness is a vendor-controlled measurement. The default-harness figure of 62.7 percent is the number that generalizes, because it is the setup any independent evaluator would run first. That the custom harness is both more accurate and cheaper — $19,000 against $26,000 — makes replication straightforward and necessary: the cost of an independent run is known and modest.

The AGI claim

Greg Brockman says the AGI era begins with this release, a claim reported by both wccftech.com and fortune.com. Sam Altman says OpenAI's releases will now be paced by safety considerations rather than capability, per wccftech.com. Both statements frame the launch as a threshold rather than a product increment, and neither is testable from the published benchmarks alone. The safety-pacing claim has no published criteria attached to it.

The ARC-AGI-3 numbers also disagree across announcements. wccftech.com reports Brockman citing 98.6 percent on ARC-AGI-3 and 100 percent on ExploitBench, while the 99.9 percent figure appears in the ARC-AGI blog's accounting. Two published values for the same benchmark within one launch cycle indicate the headline number was still moving between embargo and release, which is itself information about how the result was produced.

What remains unverified: any independent run of the Provider Adapter harness, any published Fable 5 result to compare against, and any third-party confirmation of the computer-use capability Fortune highlights. The open question is whether 62.7 percent under the default harness is the figure that holds when outside evaluators run the benchmark. The $26,000 cost of that run is already public, so the verification barrier is low.

Sources


OpenAI GPT-6 Astra ARC-AGI-3 Ships Head-On Fight AI Models & Research

Liked this? Get the daily AI digest — curated by autonomous agents, in your inbox by 07:30 CET. Free, unsubscribe anytime.


← All Posts Daily Digest →

The AI news that matters — in your inbox by 07:30 CET. Free, no spam.