注册并分享邀请链接,可获得视频播放与邀请奖励。

Rohan Paul (@rohanpaul_ai) “More good news for local LLMs. Tencent’s new Hunyuan Hy3 reaches Gemini 3.5-leve” — TopicDigg

Rohan Paul 的个人资料封面
Rohan Paul 的头像
Rohan Paul
@rohanpaul_ai
Compiling in real-time, the race towards AGI. The Largest Show on X for AI. 🗞️ Get my daily AI analysis newsletter to your email 👉
加入 June 2014
6.9K 正在关注    153.4K 粉丝
More good news for local LLMs. Tencent’s new Hunyuan Hy3 reaches Gemini 3.5-level physics quality for 35x less cost. Test was done on atomic[.]chat, a desktop app that runs LLMs locally. The prompt asked 4 models to build bowling, air hockey, and pool simulations. The harder part was preserving physical cause and effect. A strike needs collision timing, mass transfer, pin rotation, friction, and believable scattering. A pool break exposes the same weakness, because every wrong angle compounds immediately. Interestingly, DeepSeek-V4 spent the highest number of tokens (50,600 ), yet produced the weakest visual physics in this test.
显示更多
New Hunyuan Hy3 hits Gemini 3.5 quality on physics for 35x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: - A bowling ball knocking down the pins - An air hockey rally that ends in a goal - A pool break scattering the rack Outputs: Hunyuan Hy3: 29,757 tokens, $0.006 Gemini 3.5: 23,300 tokens, $0.21 GLM-5.2: 25,454 tokens, $0.07 DeepSeek-V4: 50,600 tokens, $0.009 Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenes
显示更多
0
4
60
10
转发到社区