The "cheap AMD alternative" to the $4,500 NVIDIA DGX Spark costs $3,999 fully specced. That's not an alternative. That's the same price.
Three compact machines are being sold as the answer to local AI: the DGX Spark at ~$4,500, the AMD Strix Halo at ~$4,000, and a Mac at roughly half of either. All three marketed on the same promise — run big models at home, skip the cloud.
But the benchmarks everyone shares measure the wrong thing.
They race token generation — the speed you watch text appear — and on that number all three nearly tie. The Mac generates at 56 tokens/sec. The $4,500 DGX? 84. Barely a gap you'd feel in a chat.
The number that actually separates these machines is prefill — how fast each one reads your context before it answers. And there the gap is brutal: the DGX hits 2,107 tokens/sec, six times the AMD box.
So the real question was never "which is fastest." It's "what's your bottleneck":
- Chatting and short prompts → generation matters, and the cheaper machine already wins
- Long context, RAG, agents, huge files → prefill matters, and only the DGX delivers
- Big models no consumer GPU can hold → AMD's 128GB unified memory, up to 96GB as VRAM
One honest test on the same 30B model settled all three. The full breakdown — prefill vs generation, price vs bottleneck — is pinned 👇
显示更多