这位开发者将 8 台 NVIDIA DGX SPARK 连接成一个集群 - 并运行了一个让他生产力提升 10 倍的 800GB 模型
21:47 他直言不讳地说 - “这是一个 TB 的 VRAM - 我们运行了 Quen 3.5,800GB 磁盘上,一个甚至无法装入单台 Mac Studio 的模型 - 24 个 token 每秒 - 我得说这是个胜利”
8 台 Spark 通过价值 1,300 美元的交换机通过以太网上的 RDMA 连接 - 每个节点添加 128GB 内存到一个统一的 1TB 内存池中
从一台 Spark 开始,每秒 3 个 token - 每添加一个节点速度翻倍 - 八台一起为一个物理上无法在其他地方运行的模型提供 24 个 token
Kimi K2 600GB 在 15 分钟内加载,每个节点 115GB,每秒 13 个 token - 一个根本无法在任何更小设备上运行的模型
Claude 帮助配置了整个集群 - 所有 8 台机器的 SSH 网格、网络配置、巨型帧、QSFP 端口速度 - 全部从一个终端完成
大多数人为这种规模的模型租用云端计算,每月 2,000 美元以上 - 他一次性构建了集群,现在每个 token 的成本降低了 20 倍
THIS DEVELOPER CONNECTED 8 NVIDIA DGX SPARKS INTO ONE CLUSTER - AND RAN AN 800GB MODEL THAT MADE HIM 10X MORE PRODUCTIVE
21:47 he says it straight - "this is a terabyte of VRAM - we ran Quen 3.5, 800GB on disk, a model that doesn't even fit on a single Mac Studio - 24 tokens per second - I'd say that's a win"
8 Sparks connected through a $1,300 switch via RDMA over Ethernet - each node adding 128GB of memory into one unified pool of 1TB
started with one Spark at 3 tokens per second - every added node doubled the speed - and eight together deliver 24 tokens on a model that physically cannot run anywhere else
Kimi K2 at 600GB loaded in 15 minutes, 115GB per node, 13 tokens per second - a model that simply cannot run on anything smaller
Claude helped configure the entire cluster - SSH mesh across all 8 machines, network config, jumbo frames, QSFP port speeds - all from one terminal
most people rent cloud compute for models this size at $2,000+/month - he built the cluster once and now every token costs 20x less
显示更多