An editor compares two printed AI launch pages on a whiteboard, red circles on changing hallucination and benchmark figures
A leaderboard is not stone; sometimes it is a still-refreshing press page., AI-generated illustration, not a news photograph

GPT‑6 Astra’s September 3 launch post did more than ship a model—it rewrote its own scoreboard. Fortune compared archival snapshots: the page was briefly pulled (CMS bug, then an outage, OpenAI said), came back with several figures friendlier to Astra, and kept changing. The hallucination rate dipped from about 4.2% to 2%, then returned to 4.2%; some math scores for predecessor Sol and rival Fable 5.1 also slid around.

This is not another retread of the ARC harness gap—that story is already on the record. The point here is narrower: a public scoreboard that can still be quietly edited on launch afternoon.

[1][2]

Which dials moved in the archives

Fortune and a later TNW wrap list the churn: beyond hallucination ping-pong, Sol’s ExploitBench briefly jumped from ~5.5% to 11.5%—OpenAI said it may revert because that reasoning tier is not a commercial offering—and the embargo draft’s ARC‑AGI‑3 98.6% became 99.99% live. Not every edit flattered Astra; some Anthropic HealthBench numbers rose too. The pattern still feels like storyboards yielding to launch narrative.

OpenAI told Fortune evaluations carry noise across checkpoints, scaffolds, and runs, and that updates aimed at best-available comparisons. Stanford researchers countered that the system card barely documents the internal hallucination eval—even test-item counts—and coined a blunt label: benchmaxxing.

[1][2]

Why post-hoc edits hurt more than one hot score

Buyers and investors use launch pages as procurement and valuation baselines. If numbers drift between archive snapshots, the baseline dissolves. Stack that on harness dependence—tens of points between a standard harness and a Provider Adapter—and outsiders cannot tell whether they are comparing weights, scaffolding, or copy versions.

Practitioners like Vincent Sunn Chen propose a boring norm: when you change a published figure, say what changed. Boring, and more auditable than “trust the noise.”

[1][2]

Spillover into a 2027 listing story

As Silicon Report and others note, metric drift makes it harder for customers and investors to pick models for tasks—right as an OpenAI 2027 IPO window is discussed. Rivals will weaponize archive diffs; procurement may start demanding frozen system-card versions plus third-party repros instead of blog screenshots.

[1][2]

Take

I read the snapshot chase as a trust tax. Labs may fix honest errors, but without a changelog, fix and polish look identical in pixels. The industry norm worth fighting for is not a higher headline score—it is a diff on every public revision: old value, new value, reason, and whether the commercial config changed. Skip that, and Astra’s launch page is just a marketing PDF that breathes.

[1][2]