THIS ENGINEER CONNECTED A DGX SPARK AND A MAC STUDIO OVER A 50 GIGABIT LINK AND BUILT A HYBRID AI CLUSTER THAT BEATS BOTH MACHINES RUNNING ALONE
the idea is simple - DGX Spark is insanely fast at processing prompts at up to 1,700 tokens per second but slow at generating them - Mac Studio is the opposite, slow on prefill but fast on decode at 106 tokens per second
so he split the work - Spark handles prefill, Mac Studio handles decode - two machines doing exactly what they're optimized for over a 50 gigabit network link
the result: Spark-class time to first token combined with Mac-class decode speed - neither machine achieves this alone
at 8B models the Mac Studio decoded 8x faster than the Spark - and disaggregated the setup captured both advantages simultaneously
at 32B models the Spark processed prompts 2.5x faster than the Mac Studio on prefill - and disaggregation recovered that speed while still getting faster decode
the 50 gigabit Mellanox ConnectX4 link added only 18 milliseconds of overhead - essentially zero compared to the gains
Claude Code did most of the setup - SSH into both machines, cloned repos, compiled VLLM from source, built Rust networking bindings and MLX with Metal shaders - while he watched
two machines, one 50 gig link, one hybrid cluster that runs frontier models faster than either box could alone
显示更多