Grok 4.6 just ranked #1# on CursorBench 3.2
Outperforming Claude Fable 5, Opus 5 and GPT-5.6 Sol on real-world coding performance
And what makes this even crazier is the efficiency....the chart gives CursorBench performance against average cost per task, and Grok 4.6 is sitting right at the top of the frontier
Grok 4.6 is insanely capable at coding....delivering frontier coding performance at ridiculous efficiency
Grok 4.6 from @SpaceXAI on ARC-AGI (Verified):
- ARC-AGI-1: 87.5%, $0.30/task
- ARC-AGI-2: 67.1%, $0.76/task
- ARC-AGI-3: 2.11%, $5.6K
On ARC-AGI-3, Grok 4.6 with xhigh reasoning scored comparably to GPT-5.6 Sol with high reasoning, but cost $5.6K versus Sol's $15.2K.
Ultrafast mode for GPT-5.6 Sol is now in limited preview, powered by Cerebras.
We gave @OpenAI's GPT-5.6 Sol the same prompt on Ultrafast and Standard: build a financial terminal-style dashboard for analysts.
Ultrafast: 1 min 50 seconds
Standard: 12 min 20 seconds
Same result, nearly 7x faster.
Grok 4.6 wins again. 👑
Grok 4.6 takes the #1# spot on GPQA Diamond with a score of 94.9%, beating GPT-5.6, Gemini 3.1 Pro, Claude Opus 5, and every other model tested by Artificial Analysis.
Grok 4.6 ranks #1# on the GPQA Diamond leaderboard 🧠
Grok 4.6 (high) scores 95% - the highest score on the chart for graduate-level scientific reasoning
It outperforms Claude Fable 5, Opus 5, GPT-5.6 Sol and Kimi K3
Previewing Ultrafast mode for @OpenAI's GPT 5.6 Sol, powered by Cerebras.
GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model -- up to 14× faster than the same model on Standard processing.
It speedran Humanity’s Last Exam in 11h 11m, nearly 7× faster than Claude Fable 5 with comparable accuracy.
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed.
Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.