GPT-5.6 Sol Ultra lost to GPT-5.5 on physics at 3× the cost!
We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
Prompts:
- A monster truck backflipping onto a parked car
- A stunt car jumping six buses into a brick wall
- A train derailing off a broken bridge into the water
Outputs:
GPT-5.6 Sol Ultra: 32.9K tokens, $0.33
Opus 4.8: 9.2K tokens, $0.24
GPT-5.5: 12.4K tokens, $0.11
Grok 4.5: 7.0K tokens, $0.08
Sol Ultra draws just like GPT-5.5, only with more detail and it shows clearest on the bus jump where the two look almost identical. In the other tests Sol Ultra came out worse than its predecessor and the physics really got weaker. We think GPT-5.5 took the truck flip and the train outright. With Sol Ultra you basically get GPT-5.5 with weaker physics and a nicer picture for 3x the price. Newborn Grok 4.5 failed two of the three tests and only came good on the train
显示更多
Sol, Terra, and Luna, our GPT‑5.6 family of models, are starting to roll out now in ChatGPT, Codex, and the API.
New Hunyuan Hy3 hits Gemini 3.5 quality on physics for 35x cheaper!
We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
Prompts:
- A bowling ball knocking down the pins
- An air hockey rally that ends in a goal
- A pool break scattering the rack
Outputs:
Hunyuan Hy3: 29,757 tokens, $0.006
Gemini 3.5: 23,300 tokens, $0.21
GLM-5.2: 25,454 tokens, $0.07
DeepSeek-V4: 50,600 tokens, $0.009
Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenes
显示更多
Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8!
We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
Prompts:
— A train derailing off a broken bridge into the water
— Two cars jumping off ramps and colliding mid-air over a canyon
— A monster truck crushing a row of parked cars
Outputs:
Fable 5: 62,158 tokens, $3.12
GPT 5.5: 37,753 tokens, $1.14
Opus 4.8: 22,280 tokens, $0.56
GLM 5.2: 36,246 tokens, $0.08
Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.
显示更多
New Claude Sonnet 5 performs at GPT 5.5 level 6x cheaper!
We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos
Prompts:
- A car crashes into a brick wall
- A wrecking ball destroys a house
- A catapult throws a rock at a castle wall
Outputs:
Sonnet 5: 15,047 tokens, $0.15
Opus 4.8: 23,063 tokens, $0.58
Sonnet 4.6: 25,824 tokens, $0.39
GPT 5.5: 31,152 tokens, $0.94
Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other model
显示更多
Introducing Claude Sonnet 5, our most agentic Sonnet yet.
It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models.
显示更多
Mistral OCR 4 turned a handwritten calculus exam into clean LaTeX!
We gave it a photo of a hand-written exam page. The model read the handwriting and rebuilt every formula into structured digital text
Output: Time: 5.1s · Cost: $0.09
Formulas came through exactly right - the hard part was nailed. The graph, unfortunately, it didn’t redraw. But that’s the telling part: most OCR tools just dump the text and quietly drop the figure. OCR 4 caught the plot, boxed it, and tagged it as a chart. It doesn’t get redrawn, but it gets read and accounted for
显示更多
Introducing Mistral OCR 4. It creates structure with bounding boxes, block classification, and inline confidence scores in 170 languages. 🧵👇
Sakana Fugu surprisingly performed near GLM 5.2 level but 17× more expensive!
We gave the same prompt to 4 models: build a complete live Trader Desk with both frontend and backend components, real-time market data fetched from external APIs for 8 symbols, and a custom dark-theme UI.
Outputs:
Fugu Ultra — 22,225 t, $0.51
Opus 4.8 — 15,802 t, $0.31
GPT-5.5 — 11,474 t, $0.26
GLM 5.2 — 13,677 t, $0.03
Fugu created the most polished and feature-rich trading desk in the run. GLM 5.2 was very close behind, with a similarly complete multi-panel interface and live data, but at a much lower cost. Opus and GPT also performed well, delivering solid results with a better balance between quality and cost
显示更多
Nemotron 3 Ultra performed GPT 5.5 level 10× cheaper
We gave three same prompts to build HTML5 canvas with real physics. At first scene we have water in a spinning drum. Galton board - balls through pegs into bins. And a block collision setup with extreme mass differences.
Outputs:
Nemotron 3 Ultra: 11.3k tokens, $0.051
GPT 5.5: 11.0k tokens, $0.57
Nemotron stays right on GPT 5.5's heels, but at 10× cheaper. The gap in quality is far smaller than the gap in price.
显示更多