Night newsroom desk: twin monitors show ARC-AGI-3 STANDARD 62.7% and PROVIDER ADAPTER 99.9%, with a server blade wrapped in a climbing harness between them
Illustration: same weights, two scaffolds, two scoreboards (AI-generated, not a news photo), AI-generated illustration, not a news photograph

When OpenAI shipped GPT-6 Astra, the number that traveled was 99.9% on ARC-AGI-3. Three days later, the number that should stay on the wall is the other one: 62.7%.

Both come from the people who built the ruler. ARC Prize’s September 3 post lays out the dual runs: the Standard harness at max reasoning scored 62.7% for about $26,098; the Provider Adapter at high reasoning scored 99.9% for about $18,817. The prettier score was also the cheaper run.

The gap is not the weights. It is the scaffold. The standard line gives every model the same minimal interface and lets the model decide which notes to keep. OpenAI’s adapter preserves opaque reasoning state between calls and compacts long conversations—vendor context management welded into the exam. The row that should make marketers flinch: set reasoning effort to none inside the adapter and Astra still scores 96.7%, more than thirty points above the same model at max effort on the standard line. The harness beat the reasoning dial.

The viral comparison was rougher still. Feeds compared 99.9% with GPT-5.6 Sol’s 7.8%. TNW notes those are not the same test: adapter versus standard. Like-for-like is 62.7% versus 7.8%—still a leap, no longer an AGI ticket. ARC Prize refused the slogan outright (“not claiming that it is AGI”); Mike Knoop wrote that evidence is still missing. Going forward, both harnesses will sit side by side on the leaderboard.

Then the published numbers moved. TNW cites Fortune’s archive diffs: at least five metrics on the launch page shifted after publication—Astra’s hallucination rate bouncing between 4.2% and 2%; Fable 5.1’s FrontierMath sliding from 87.8% to 78% before settling at 83%; Sol’s ExploitBench doubling from 5.5% to 11.5% (OpenAI said it may revert; that figure reflects a reasoning tier not sold commercially). The embargo draft had ARC-AGI-3 at 98.6%; the live post said 99.99%. Stanford researchers Anka Reuel and Mike Hardy call re-running until the number improves benchmaxxing. On the internal hallucination benchmark, they found “barely any details” in the system card.

None of that erases real progress. 62.7% on the standard line is a step-function against near-zero early scores, and adapter runs beat the human median on action efficiency. But the industry increasingly sells “the model” and “the productized scaffold around it” as one marketing noun—Anthropic, Google, and Microsoft price harnesses; Nvidia’s scaffolding took Claude Opus 5 from 30.2% to a clear. You buy weights. The screenshot sells a system.

Independent indexes kept the romance short. On Artificial Analysis, Astra’s Intelligence Index is 61—tied with the model it replaces, behind Fable 5.1—and its Coding Agent Index sits at 67 under Fable 5.1’s 70. Hallucination and long-horizon knowledge work improved; economically weighted tasks, banking support, scientific Python, and long-context reasoning slid. The louder the slogan, the more the evaluation condition belongs in the same sentence.

Take: Astra is worth serious use, especially for computer-use and engineering agents. “Welcome to the AGI era,” if it rests on misaligned comparison tables, is stage audio—not a scientific finding. Next time a near-perfect score goes viral, ask three questions first: what is the standard-line number, what did the adapter do, and did the comparison model wear the same rope.

[1][2]