Mac Mini on Metal: 55.64 t/s.
NVIDIA rig, two cards: 83.65 t/s.
Same model. Same quant. Guess where that gap actually comes from.
Qwen3-Coder 30B (A3B active), Q4_K_M on all three boxes. The Mac isn't dying. It's ~33% behind a real GPU stack on token generation. On a chip that fits in a lunchbox.
Where NVIDIA runs away is prefill. pp512: 2107 vs 563 t/s. That's raw compute. Prompt processing is compute-bound.
Token generation isn't. tg is bandwidth-bound. M-series unified memory sits around 400 GB/s. RTX 3090 sits at 936. That's your ratio. tg128 confirms it almost line for line.
Which is the whole reason a used 3090 stomps a $4k DGX Spark at 273 GB/s. And why the guys dropping $2–4k on Ryzen AI Max+ 395 keep asking where the tokens went.
ngl the "best rig" everyone chases is the wrong axis. What's actually in your box?
显示更多