注册并分享邀请链接,可获得视频播放与邀请奖励。

Bleys Goodson (@bleysg) “It is a 2 to 4T param model. They are serving it across 70-100 wafers. To get he” — TopicDigg

Bleys Goodson 的个人资料封面
Bleys Goodson 的头像
Bleys Goodson
@bleysg
Helping people engineer the future.
加入 March 2009
1.4K 正在关注    1.2K 粉丝
It is a 2 to 4T param model. They are serving it across 70-100 wafers. To get healthy serving characteristics, they are essentially putting at most one layer per wafer, and the model is in the ballpark of 70-90 layers. There's a couple of different ways this could be served and model sizes implied by that. One is if they keep the heavy KV caches they've used before. Another is if they go with lighter KV cache designs more akin to DeepSeekV4 or Hybrid SSM models. The fact that they've partnered with Cerebras and designed with the hardware in mind means they're much more likely to have gone the second route. That SRAM bandwidth is too precious for a heavy KV cache. As such, something like the below is the center of probability mass: 3T total, 150B active, 70 layers.
显示更多
0
15
719
65
转发到社区