Benchmarks

Grok 4.6: 61 on the Intelligence Index, Frontier Again

SpaceXAI shipped Grok 4.6 on August 12: 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol, with pricing unchanged at $2 and $6 per million tokens.

Grok 4.6: 61 on the Intelligence Index, Frontier Again — article cover

On August 12, 2026, SpaceXAI released Grok 4.6, just over a month after Grok 4.5 — a fast cadence for a frontier lab. The same day, Artificial Analysis published its independent evaluation: a score of 61 on the Intelligence Index, up 5 points from Grok 4.5 and 23 points above Grok 4.3, matching GPT-5.6 Sol (max), trailing only Claude Opus 5 (max, 63) and Claude Fable 5 (62), and just ahead of Kimi K3. The official announcement and the benchmark analysis together drew close to a thousand combined points on Hacker News. One modest, fast-cycle update, and SpaceXAI is back in the frontier group.

The Score: Back at the Frontier

What does 61 mean? It sits on the same tier as GPT-5.6 Sol (max), one to two points behind Claude’s two flagships — a tighter top of the index than at any point this year. For context, the effective frontier on this benchmark is now a cluster of models rather than a single leader. On sub-evaluations, GDPval-AA v2 Elo reached 1753, behind only Claude Opus 5 and statistically tied with Claude Fable 5 and Qwen3.8 Max. It is also worth reading how the two scorecards differ. x.ai’s own published eval table claims leads on GDPval-AA, CursorBench v3.2, AA-Briefcase, and Harvey LAB, but lists DeepSWE behind GPT-5.6 Sol Max (65.9% vs 73%) and trails FrontierCode and APEX-SWE against Fable 5 Max. A vendor acknowledging weak spots inside its own launch table remains rare among frontier releases in 2026, and it makes the third-party numbers easier to trust.

Agentic Work Is the Real Battleground

Artificial Analysis puts Grok 4.6’s strength plainly: few models are simultaneously competitive across knowledge work, customer service, and terminal tasks, and this one “sits on the cost-vs-performance Pareto frontier for every agentic evaluation.” The sub-scores back that up: on tau3-Banking it scored 50.7%, in the top two alongside Qwen3.8 Max at 51.3%, and its AA-Briefcase Elo of 1577 lands in Claude Fable 5 territory. The deeper story is turn efficiency. Completing the index tasks took about 53 turns and roughly 0.5 billion accumulated input tokens on average, versus about 103 turns and 2.0 billion tokens for Claude Opus 5 (max). Agent costs are driven by turns and context accumulation rather than sticker price, so a model that finishes the same work in half the turns can deliver bills far below what its per-token pricing suggests — and that is exactly where Grok 4.6’s flat pricing compounds.

Why Unchanged Pricing Matters

Pricing stays at Grok 4.5 levels: $2 per million input tokens and $6 per million output tokens, with cache hits nudged from $0.30 to $0.50 per million and the context window holding at 500k tokens. Artificial Analysis estimates the measured cost per task at $0.84 — the same as Kimi K3 — and notes Grok 4.6 is more than 60% cheaper than Claude Opus 5 ($5/$25) or GPT-5.6 Sol ($5/$30). A frontier model shipping a full generation of intelligence gains without a price increase is unusual in 2026; gains at the frontier normally come with hikes. Because output token price dominates the cost of reasoning-heavy workloads, the $6 output rate is doing quiet, heavy lifting here, and it flows straight into model-selection decisions for any team running agents at volume. Kimi K3 matching that per-task cost is notable in its own right: the open-weight camp is now setting the price reference that frontier labs have to beat.

Availability and Developer Entry Points

Grok 4.6 is available on day one in Cursor and Grok Build, via the API with keys from console.x.ai and documentation at docs.x.ai, and through partners including OpenRouter, Vercel, and Cloudflare. Technically, x.ai describes a longer supplemental training run, SFT trajectories regenerated from Grok 4.5 with model-based filtering, and agentic reinforcement learning across coding, knowledge work, kernel optimization, web development, and CAD. The company says the model is better at “checking its own work before moving on,” with improved one-pass results on visual and interactive projects, and that it ran its “widest-ever suite of pre-deployment testing” plus third-party and post-deployment evaluation. A doubled-usage promotion runs for the first week in Grok Build and Cursor, and the “fast” variant costs double the standard rates.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL