注册并分享邀请链接,可获得视频播放与邀请奖励。

Rulya (@Rulyaxd) “Mac Mini on Metal: 55.64 t/s. NVIDIA rig, two cards: 83.65 t/s. Same model. Same” — TopicDigg

Rulya 的个人资料封面
Rulya 的头像
Rulya
@Rulyaxd
never look back
加入 April 2022
801 正在关注    151 粉丝
Mac Mini on Metal: 55.64 t/s. NVIDIA rig, two cards: 83.65 t/s. Same model. Same quant. Guess where that gap actually comes from. Qwen3-Coder 30B (A3B active), Q4_K_M on all three boxes. The Mac isn't dying. It's ~33% behind a real GPU stack on token generation. On a chip that fits in a lunchbox. Where NVIDIA runs away is prefill. pp512: 2107 vs 563 t/s. That's raw compute. Prompt processing is compute-bound. Token generation isn't. tg is bandwidth-bound. M-series unified memory sits around 400 GB/s. RTX 3090 sits at 936. That's your ratio. tg128 confirms it almost line for line. Which is the whole reason a used 3090 stomps a $4k DGX Spark at 273 GB/s. And why the guys dropping $2–4k on Ryzen AI Max+ 395 keep asking where the tokens went. ngl the "best rig" everyone chases is the wrong axis. What's actually in your box?
显示更多