On Agents' Last Exam, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive) by 13.1 points.
At medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. GPT‑5.6 Terra and Luna also outperforms Fable 5 at around one-sixteenth the cost.
显示更多