BREAKING: Grok 4.5 leads VulcanBench’s new coding benchmark. 🔥
Grok scored 91.3%, solving 21 of 23 real-world software tasks across five languages, beating Claude Fable 5 and GPT-5.6 Sol while also owning the cost-efficiency frontier.
Grok keeps winning. 🏆
显示更多