注册并分享邀请链接,可获得视频播放与邀请奖励。

Feiteng (@FeitengLi) “智谱 GLM-5.2 开源模型,干到全球前三、持平 Claude Opus 4.8,长程编码还超 GPT-5.5” — TopicDigg

Feiteng 的个人资料封面
Feiteng 的头像
Feiteng
@FeitengLi
Speech · Image · Video · LLM multimodal generative models Research → Infrastructure → Built AI algo serving 200M+ users 公众号: Generative AI Open to opportunities
加入 November 2016
1.3K 正在关注    3.5K 粉丝
智谱 GLM-5.2 开源模型,干到全球前三、持平 Claude Opus 4.8,长程编码还超 GPT-5.5、成本只要 1/6。 但 5.0→5.1→5.2 这三代到底改了啥,挺容易搞混: 5.1 没动架构,纯后训练刷代码(48.3→68.7)。 5.2 是真大改:上下文 200K→1M,IndexShare 每 4 层共享一个 indexer,1M 下 per-token FLOPs 砍 2.9×;MTP 升级,接受长度 +20%。
显示更多
(Claude、GPT、GLM) GLM-5.2 Tops Artificial Analysis as the #1# Open-Source Model, Ranking Top 3 Globally GLM-5.2 launched and went open-source today, delivering a solid scorecard across multiple authoritative third-party benchmarks and arenas. 📊 Artificial Analysis Intelligence Index A comprehensive evaluation that integrates several authoritative leaderboards spanning coding, reasoning, long context, and more. GLM-5.2 scored 51, ranking among the top of all available models—on par with Claude Opus 4.8—and claiming the #1# spot among open-source models worldwide. 🎨 Code Arena A real-world head-to-head arena focused on front-end code generation, with Elo rankings produced by blind user voting. GLM-5.2 ranked #2# globally with a score of 1,595. 🏆 DesignArena A category arena centered on scenarios that combine design and code. GLM-5.2 took the top spot with a score of 1,360. ⚙️ FrontierSWE A software-engineering benchmark built around the "frontier of human capability," assessing engineering ability across three dimensions: implementation, performance, and research. GLM-5.2 ranked #3# overall. 💪 From front-end development and design-to-code to engineering-grade software tasks, GLM-5.2 consistently lands in the top tier across multiple real-world evaluation scenarios, steadily closing in on the world's strongest models. We'll keep pushing forward in pursuit of an ever-higher ceiling of intelligence.
显示更多
0
22
43
2
转发到社区