Grok 4.5 is the top non-Anthropic model on AA-Briefcase, combining frontier agentic knowledge work capabilities with leading cost and time-efficiency
Yesterday
@SpaceXAI released Grok 4.5, a new frontier-level model with strengths in agentic coding and knowledge work. On AA-Briefcase, Grok 4.5 scores 1328, a +578 improvement over Grok 4.3 and the highest score of any non-Anthropic model (note that GPT-5.6 not released yet). It achieves this while sitting on the cost and time efficiency frontier, averaging $1.12 per task, 86% lower than Claude Opus 4.8 (max), and 12.4 minutes per task, around half the time of Opus 4.8 (max).
AA-Briefcase is our new proprietary benchmark for agentic knowledge work, testing models on a fully private dataset of realistic tasks across thousands of complex input files. Tasks require deliverables like spreadsheets, presentations, and UI mock-ups, with performance combined into a single AA-Briefcase Elo across correctness, analytical quality, and presentation quality.
Key results for Grok 4.5 with high reasoning on AA-Briefcase:
➤ Frontier agentic knowledge work capabilities: Grok 4.5 achieves an AA-Briefcase Elo of 1328, the highest score of any non-Anthropic model, behind only Claude Fable 5 (1390), Claude Sonnet 5 (max, 1390), and Claude Opus 4.8 (max, 1354). Across the three AA-Briefcase scoring axes, Grok 4.5 is strongest on objective rubric criteria and analytical quality, with comparatively weaker presentation quality. It achieves the second-highest overall rubric pass rate (40.7%), behind Claude Fable 5 (56%) and Claude Sonnet 5 (42.3%)
➤ Leading cost efficiency: Grok 4.5 averages a cost of $1.12 per AA-Briefcase task, placing it on the cost-performance Pareto frontier. This is much most cost effective than peer models such as Claude Opus 4.8 (max, $8.26) and GLM 5.2 (max, $1.71)
➤ Faster task completion: Grok 4.5 averages 12.4 minutes per AA-Briefcase task, also placing it on the speed-performance frontier. It is much faster than Claude Opus 4.8 (max, 23.9 min) and Claude Sonnet 5 (max, 36.9 min), primarily due to lower turn use. Grok 4.5 averages just 23 turns per task, ~40% of GLM 5.2 (max, 56) and ~13% of Claude Sonnet 5 (max, 183)
Congratulations to
@SpaceXAI,
@cursor_ai, and
@elonmusk on the impressive release!