注册并分享邀请链接,可获得视频播放与邀请奖励。

Aaron Burnett (@aaronburnett) “Grok 4.5 is looking like a success with help from Cursor data but underneath the” — TopicDigg

Aaron Burnett 的个人资料封面
Aaron Burnett 的头像
Aaron Burnett
@aaronburnett
Founder, CEO @mach33 | Research and Investment in Expansion Technologies
加入 July 2011
607 正在关注    24.9K 粉丝
Grok 4.5 is looking like a success with help from Cursor data but underneath the surface we expect future Grok/Cursor model training is likely to speed up in the coming months. We've spent several months getting up to speed on the SpaceXAI business, especially the tech underneath it. The C rewrite was under appreciated by the investor community so we dug in to quantify its impact. including building a physics first model that functions as a stopwatch for the SpaceXAI model factory. Bottomline: C-rewrite gets SpaceXAI faster model cycles, leveraging 33% more tokens/second/GPU against SOTA competition resulting in the potential to shipping new models every ~3.5 weeks. two core learnings from this modeling exercise: 1: the training cycle speed up is primarily coming from RL (not pretraining) where the increased tokens/second/GPU advantage can shave up 2+ weeks off full model training cycle. 2: the rewrite itself should compound the time savings as model sizes grow. at ~2T shaving off 2-3 weeks, ~8 weeks at 6T, and 15 weeks at 10T. Note: a 20T parameter model likely runs into a data bottleneck prior to a training speed bottleneck but the directional advantage stands. Also we assume tokens/second/gpu advantage will melt over time as competitors try to match it. when you do the math, in true SpaceX and Elon fashion, it looks like they are attempting to build a SOTA model factory that can pump out bigger models faster than anyone else. Full analysis here for the public:
显示更多
0
12
519
50
转发到社区