On Monday afternoon, September 21, xAI — now operating under the SpaceXAI brand — released Grok 4.7, officially positioned as "our most capable Grok yet, built for coding and knowledge work." Pricing is identical to Grok 4.6: $2 per million input tokens, $6 per million output. But the underlying base model is larger than 4.6, with a substantially extended reinforcement learning cycle — the battleground this generation is explicitly set on is long-horizon tasks.
[1][2]The nut graf: Grok 4.7 is a machine that keeps the sticker price and lands mid-pack. On coding and knowledge-work benchmarks it nearly overtakes GPT-5.6 Sol Max and chases Fable 5.1 Max, yet on the Artificial Analysis Intelligence Index it scores 46, a tier below Claude Fable 5.1 and GPT-6 at 53. Same-price upgrade plus mid-tier breakout is the correct coordinate for reading this generation: xAI chose to convert a larger base model and a longer RL cycle into "long-task capability at no extra cost," rather than chasing parity on all-around leaderboards.
Several numbers deserve unpacking. Decrypt reports Grok 4.7 runs on roughly 2.1 trillion parameters, with supplemental training on SpaceX engineering data — a proprietary data asset that is xAI's most distinctive divergence from competitors trained on internet text alone. On specific benchmarks, as announced by xAI and relayed by media, DeepSWE stands at 71.0%, CursorBench 4.0 at 46.3%, Terminal-Bench 38.0%, and EEBench 66.0% — all inside the knowledge-work-and-coding cluster. Musk's self-assessment on X was that the model achieves "a very competitive balance between intelligence, speed, and cost," adding that Grok 4.8 has already finished training — one post that simultaneously defends the pricing and previews the next generation.
The disappointment criticism is also fair. 36Kr's coverage literally headlines "price war, disappointing scores": on the broadest metric, 46 points parks it mid-tier of the first group, without shaking the Claude and GPT-6 headship. The long-task narrative is precisely a play to strengths — a model that wastes fewer steps across multi-hour tasks and returns directly usable results is closer to the developer market xAI wants than single-question IQ tests. The problem: long-horizon capability currently lacks a widely accepted standardized metric like the AA Index, so "best at long tasks" is hard to verify by third parties — the sharpest dispute around this generation.
Put Grok 4.7 in the pricing context of 2026 and the strategy sharpens further. Frontier API pricing has been sliding all year — cache read prices, flash tiers, and open-weight alternatives at fractions of flagship rates — while OpenAI and Anthropic both hold their $10/$50 flagship price points. xAI's $2/$6 for a model this size is not a discount play; it is an anchor. It tells developers that agentic workloads, which burn tokens across multi-hour runs, should be priced like a utility, not a premium. The Decrypt headline's complaint that xAI is "late to the party" captures the timing problem: Grok 4.7 arrives after Claude Fable 5.1 and GPT-6 have already defined the flagship conversation, so xAI's only credible wedge is the price-to-long-horizon ratio. Whether that wedge holds depends on the ecosystem xAI can build around it — IDE plugins, agent frameworks, and enterprise deployment paths — because model quality alone no longer differentiates at this tier. The long-task framing is coherent, but coherent framing is not yet a benchmark.
[1][2]