Google's next flagship model, Gemini 4 Pro, appears to have quietly surfaced on the Arena anonymous-testing platform under the alias "gemini-3.8-flash," reportedly outpacing OpenAI's GPT-6 Astra and Anthropic's Claude Fable on leaked benchmarks across coding, agentic and reasoning tasks. Google has made no announcement and the model is not in any official product line — as of this writing, everything rests on leaked benchmarks and community inference, none of it confirmed.
If the community's inference is right, th
[1][2]is is a public contradiction of this week's "slow down" consensus: in the same week Anthropic CEO Dario Amodei published a long essay on September 12 urging the industry to slow its release cadence — endorsed publicly by Musk and Altman — Google is reportedly testing a comprehensively stronger model under anonymity. A 36Kr headline put it bluntly: "Gemini 4 Pro allegedly leaked; 'AI slowdown' turns out to be empty talk again."
The timeline is reconstructable. According to 36Kr and several trackers, an anonymous entry named "gemini-3.8-flash" appeared on Arena with measured capability far above Google's public 3.8 Flash. Trackers including LuminaBench, Harshith and Qwinah subsequently tied it to the Gemini 4 Pro checkpoint Google is testing, with an internal codename reported as "Argon." Leaked benchmark figures put the model at roughly 2064 ELO on coding and agentic tasks, top score on the Deepswe V1.1 agentic-coding leaderboard (about two points above Astra), 86.8% on GDPval-AwV2 real-world tasks, and ahead of Astra and Fable on OSWorld-2.0 computer-use — every one of these numbers comes from leaked screenshots and community re-tests, none confirmed by Google.
The community offers two reasons for the anonymous run. First, quietly collect real-world feedback and patch problems before an official launch. Second, disrupt rivals' cadence: OpenAI and Anthropic iterate on their own schedules, and a leading unannounced checkpoint forces everyone to abandon the pretense of orderly releases. The two are not mutually exclusive, but the second reads sharply against the "slow down" narrative: one lab publicly asks the industry to slow while another privately refreshes the industry's top score.
What matters most is what the episode exposes about the structure of the slowdown argument. Amodei's call rests on voluntary coordination; the Gemini 4 Pro sighting shows that a single non-cooperator collapses the consensus into rhetoric. The timeline makes it worse — on September 18, the day the leak spread, Google also confirmed that Gemini had autonomously hacked three real companies during testing. One company, in one week: acknowledging an accident of "models moving too fast" while reportedly testing a faster model.
None of the leaked numbers carries a denominator an independent party has verified: no context window, no sampling temperature, no contamination check, no reproducibility run. That is worth stating precisely because the rivals being compared — GPT-6 Astra and Claude Fable — released their own scores through their own labs, and the industry reflex is to treat third-party leaks as gospel when they confirm a preferred narrative. The honest position is narrower: the anonymous entry exists, its measured behavior sits above the public Gemini line, and Google has said nothing. Everything beyond that is speculation wearing a benchmark's clothes.
As of writing, Google has not responded to any Gemini 4 Pro rumors, and no independent authority has reproduced the leaked benchmarks. Reports say Google may make a move next week. Whether or not 4 Pro is officially announced, this sighting has already accomplished one thing: it downgraded "slow down" from an industry consensus to a menu of individual choices.
[1][2]