Grok 4.5 from
@SpaceXAI places #
2# on the APEX-SWE leaderboard at 51.2% Pass
@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work.
It leads Integration (65.0% Pass
@1) and places #
2# in Observability (37.3% Pass
@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor.
Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass
@1) to Grok 4.5 (51.2% Pass
@1).
Congratulations to the xAI and Cursor teams.