Everyone says the DGX Spark is "2-5x faster" than the $2,000 AMD box. I benchmarked both. It's 6x on prompt processing — and near-identical on generation. Which means most people are about to overpay by $2,700.
Same 30B model, both machines:
Prompt processing (reading your context):
- AMD Strix Halo: 342 t/s
- DGX Spark: 2,107 t/s → 6x
Token generation (writing the answer, the speed you feel):
- AMD Strix Halo: 73 t/s
- DGX Spark: 84 t/s → basically a tie
Read that twice. On the number you actually watch happen — text appearing on screen — a $2,000 box ties a $4,700 one. The Spark's whole premium lives in one place: prefill.
So the buy is simple once you know your bottleneck:
- You chat, you draft, you code with short prompts → Strix Halo. Same feel, pocket $2,700, and it runs Windows and games on the side.
- Your work is huge context — long agent runs, RAG over hundreds of docs, giant files → that 6x prefill is the Spark earning its price. Nothing else touches it.
One tax the AMD hype skips: 128GB is glorious until a tool ships CUDA-only and you're debugging ROCm at midnight. The memory is real. The software lottery is too.
Buy for your bottleneck, not the biggest number on the box.
(Full breakdown of every local AI machine by bottleneck — Spark, Mac, used GPUs, this — pinned.)
显示更多