More good news for local LLMs.
Tencent’s new Hunyuan Hy3 reaches Gemini 3.5-level physics quality for 35x less cost.
Test was done on atomic[.]chat, a desktop app that runs LLMs locally.
The prompt asked 4 models to build bowling, air hockey, and pool simulations. The harder part was preserving physical cause and effect.
A strike needs collision timing, mass transfer, pin rotation, friction, and believable scattering.
A pool break exposes the same weakness, because every wrong angle compounds immediately.
Interestingly, DeepSeek-V4 spent the highest number of tokens (50,600 ), yet produced the weakest visual physics in this test.
显示更多
New Hunyuan Hy3 hits Gemini 3.5 quality on physics for 35x cheaper!
We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
Prompts:
- A bowling ball knocking down the pins
- An air hockey rally that ends in a goal
- A pool break scattering the rack
Outputs:
Hunyuan Hy3: 29,757 tokens, $0.006
Gemini 3.5: 23,300 tokens, $0.21
GLM-5.2: 25,454 tokens, $0.07
DeepSeek-V4: 50,600 tokens, $0.009
Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenes
显示更多