注册并分享邀请链接,可获得视频播放与邀请奖励。

Soran (@Soranlan) “https://t.co/2Rr5sTfntR 这位开发者将 8 台 NVIDIA DGX SPARK 连接成一个集群 - 并运” — TopicDigg

Soran 的个人资料封面
Soran 的头像
Soran
@Soranlan
拆普通人也跑得起的 AI 内容工作流 AI 热点、工具实测、生图视频、Prompt/SOP 只写跑通后的方法
加入 December 2018
335 正在关注    5.3K 粉丝
这位开发者将 8 台 NVIDIA DGX SPARK 连接成一个集群 - 并运行了一个让他生产力提升 10 倍的 800GB 模型 21:47 他直言不讳地说 - “这是一个 TB 的 VRAM - 我们运行了 Quen 3.5,800GB 磁盘上,一个甚至无法装入单台 Mac Studio 的模型 - 24 个 token 每秒 - 我得说这是个胜利” 8 台 Spark 通过价值 1,300 美元的交换机通过以太网上的 RDMA 连接 - 每个节点添加 128GB 内存到一个统一的 1TB 内存池中 从一台 Spark 开始,每秒 3 个 token - 每添加一个节点速度翻倍 - 八台一起为一个物理上无法在其他地方运行的模型提供 24 个 token Kimi K2 600GB 在 15 分钟内加载,每个节点 115GB,每秒 13 个 token - 一个根本无法在任何更小设备上运行的模型 Claude 帮助配置了整个集群 - 所有 8 台机器的 SSH 网格、网络配置、巨型帧、QSFP 端口速度 - 全部从一个终端完成 大多数人为这种规模的模型租用云端计算,每月 2,000 美元以上 - 他一次性构建了集群,现在每个 token 的成本降低了 20 倍
显示更多
THIS DEVELOPER CONNECTED 8 NVIDIA DGX SPARKS INTO ONE CLUSTER - AND RAN AN 800GB MODEL THAT MADE HIM 10X MORE PRODUCTIVE 21:47 he says it straight - "this is a terabyte of VRAM - we ran Quen 3.5, 800GB on disk, a model that doesn't even fit on a single Mac Studio - 24 tokens per second - I'd say that's a win" 8 Sparks connected through a $1,300 switch via RDMA over Ethernet - each node adding 128GB of memory into one unified pool of 1TB started with one Spark at 3 tokens per second - every added node doubled the speed - and eight together deliver 24 tokens on a model that physically cannot run anywhere else Kimi K2 at 600GB loaded in 15 minutes, 115GB per node, 13 tokens per second - a model that simply cannot run on anything smaller Claude helped configure the entire cluster - SSH mesh across all 8 machines, network config, jumbo frames, QSFP port speeds - all from one terminal most people rent cloud compute for models this size at $2,000+/month - he built the cluster once and now every token costs 20x less
显示更多