注册并分享邀请链接,可获得视频播放与邀请奖励。

vitalik.eth (@VitalikButerin) “Impressive work! For comparison, an H100 can do roughly 100-200 tok/s of Muse 30” — TopicDigg

vitalik.eth 的个人资料封面
vitalik.eth 的头像
vitalik.eth
@VitalikButerin
加入 May 2011
0 正在关注    0 粉丝
Impressive work! For comparison, an H100 can do roughly 100-200 tok/s of Muse 30B for raw inference single-thread, going up to low thousands of tok/s with a large number of threads - and I am sure that for the massively-multi-threaded case they can optimize the prover further. So we roughly, sort of, have single-digit (<10x) overhead for LLM proving! Next step is getting single-digit overheads for FHE, and then ultimately vFHE (aka STARK * FHE). A crazy ambitious milestone given present FHE overheads, but because of how highly structured and almost-linear LLM inference is, it's closer to the realm of possibility than you might think. Single-digit-overhead all the things.
显示更多
0
25
42
5
转发到社区