New Claude Sonnet 5 performs at GPT 5.5 level 6x cheaper!
We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos
Prompts:
- A car crashes into a brick wall
- A wrecking ball destroys a house
- A catapult throws a rock at a castle wall
Outputs:
Sonnet 5: 15,047 tokens, $0.15
Opus 4.8: 23,063 tokens, $0.58
Sonnet 4.6: 25,824 tokens, $0.39
GPT 5.5: 31,152 tokens, $0.94
Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other model
显示更多
Introducing Claude Sonnet 5, our most agentic Sonnet yet.
It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models.
显示更多