注册并分享邀请链接,可获得视频播放与邀请奖励。

h100envy (@h100envy) “Ex-Berkeley PhD who leads SGLang at xAI explained how they serve Grok on 100K GP” — TopicDigg

h100envy 的个人资料封面
h100envy 的头像
h100envy
@h100envy
you're literally copium, i'm your opium
加入 March 2026
34 正在关注    2.7K 粉丝
Ex-Berkeley PhD who leads SGLang at xAI explained how they serve Grok on 100K GPUs in 23 minutes - better than $2000 inference-at-scale courses. split prefill and decode -> shard experts across GPUs -> route tokens per expert -> overlap comm and compute -> serve at DeepSeek-API-killing prices. That loop is why xAI runs Grok on SGLang and third parties beat DeepSeek's own API by 5x on cost. SGLang + prefill-decode disaggregation + expert parallelism + AMD MI300 - that's the stack. Watch and save it, then read the article below.
显示更多
0
2
832
86
转发到社区