注册并分享邀请链接,可获得视频播放与邀请奖励。

Cortex (@0xCortexl) “THIS ENGINEER CONNECTED A DGX SPARK AND A MAC STUDIO OVER A 50 GIGABIT LINK AND” — TopicDigg

Cortex 的个人资料封面
Cortex 的头像
Cortex
@0xCortexl
find alpha in 2 clicks, finished school
加入 June 2023
69 正在关注    1.9K 粉丝
THIS ENGINEER CONNECTED A DGX SPARK AND A MAC STUDIO OVER A 50 GIGABIT LINK AND BUILT A HYBRID AI CLUSTER THAT BEATS BOTH MACHINES RUNNING ALONE the idea is simple - DGX Spark is insanely fast at processing prompts at up to 1,700 tokens per second but slow at generating them - Mac Studio is the opposite, slow on prefill but fast on decode at 106 tokens per second so he split the work - Spark handles prefill, Mac Studio handles decode - two machines doing exactly what they're optimized for over a 50 gigabit network link the result: Spark-class time to first token combined with Mac-class decode speed - neither machine achieves this alone at 8B models the Mac Studio decoded 8x faster than the Spark - and disaggregated the setup captured both advantages simultaneously at 32B models the Spark processed prompts 2.5x faster than the Mac Studio on prefill - and disaggregation recovered that speed while still getting faster decode the 50 gigabit Mellanox ConnectX4 link added only 18 milliseconds of overhead - essentially zero compared to the gains Claude Code did most of the setup - SSH into both machines, cloned repos, compiled VLLM from source, built Rust networking bindings and MLX with Metal shaders - while he watched two machines, one 50 gig link, one hybrid cluster that runs frontier models faster than either box could alone
显示更多