注册并分享邀请链接,可获得视频播放与邀请奖励。

NIK (@ns123abc) “GPT 5.6 SoL is a 2.2T parameter model 💀💀” — TopicDigg

NIK 的个人资料封面
NIK 的头像
NIK
@ns123abc
加入 January 2019
0 正在关注    0 粉丝
GPT 5.6 SoL is a 2.2T parameter model 💀💀
It is a 2 to 4T param model. They are serving it across 70-100 wafers. To get healthy serving characteristics, they are essentially putting at most one layer per wafer, and the model is in the ballpark of 70-90 layers. There's a couple of different ways this could be served and model sizes implied by that. One is if they keep the heavy KV caches they've used before. Another is if they go with lighter KV cache designs more akin to DeepSeekV4 or Hybrid SSM models. The fact that they've partnered with Cerebras and designed with the hardware in mind means they're much more likely to have gone the second route. That SRAM bandwidth is too precious for a heavy KV cache. As such, something like the below is the center of probability mass: 3T total, 150B active, 70 layers.
显示更多
0
19
907
19
转发到社区