NVIDIA built GPUs to do everything.
Etched thinks that's exactly the problem.
Instead of making another general-purpose chip, they built Sohu to do one thing only: run Transformer models.
No graphics. No flexibility. No legacy hardware.
Just Transformers.
That's why Etched claims Sohu delivers up to 20x the performance of an NVIDIA H100 on its target workloads.
Of course, there's a catch.
If AI moves beyond Transformers, Sohu loses its biggest advantage overnight.
It's one of the biggest bets anyone has made in AI hardware.
显示更多
Four Intel Arc Pro B70s. 128GB of VRAM total. $4,000.
One RTX PRO 6000 Blackwell. 96GB. $10,000.
The pitch writes itself: gang up cheap cards, beat the expensive one, pocket $6,000. And on raw memory it's true — four B70s hold more than the single Blackwell.
Here's what the price comparison leaves out. Four cards means the model gets split across four of them, and every token crosses PCIe to move between cards. The Blackwell is one pool — no splitting, no PCIe tax. Same reason a 128GB cluster and 128GB unified aren't the same 128GB.
So the real question isn't "which has more VRAM for less." It's what you're running. Models that fit on one B70's 32GB? The cluster's a steal. Models that need to span all four? You're now paying in latency what you saved in cash.
$4,000 of Intel is the right call for a lot of workloads. Just not because it "beats" a $10k card — because it's a different shape of machine for a different job.
显示更多
One card has 96GB of VRAM AND 1792 GB/s of bandwidth. It beats every box in this comparison on both numbers at once.
It also costs $10,000-12,000 — more than the DGX Spark, the Strix Halo, and the Mac combined, with room left over.
Now here's the part that breaks people: the $4,500 DGX Spark runs at 273 GB/s. Six times less bandwidth than this card. Same number as a Mac mini. Less than a 5090 that costs half as much.
And it still reads your context 6x faster than either of them.
Because bandwidth isn't what makes that happen. Bandwidth is how fast tokens come out. Reading your prompt in the first place is raw compute — a number that never makes it onto a spec sheet, a box, or a comparison chart.
So the DGX sits near the bottom of every bandwidth list published this year and quietly beats the machines above it at the thing most local AI work is actually waiting on.
That's why the spec sheet is the worst way to buy this hardware. Capacity tells you what loads. Bandwidth tells you how fast it writes. Neither tells you how long you'll stare at a blank screen before the first word shows up.
The 6000 Blackwell wins the chart. It also prices itself like a used car.
Full 3-way benchmark — DGX Spark vs AMD Strix Halo vs Mac, what each one actually wins — pinned 👇
显示更多
"32GB for $1,299, put four in a Threadripper, and you've got 128GB." Nice math. The card costs $1,900.
That's the listed price right now, not a hypothetical. So four of them is $7,600 — before the Threadripper board and CPU, which push the real build past $10,000.
A DGX Spark gives you 128GB unified at $4,500. An AMD Strix Halo gives you the same 128GB at $4,000.
And "128GB" means something different in each case. Four cards is 32GB × 4, split across PCIe — you're sharding models between them. The DGX and the Halo are one pool. Same number on the spec sheet, completely different machine underneath.
What the four-card build actually buys you: real bandwidth and compute per card, and a path the unified boxes can't match. What it costs: double the money, ROCm, and a build that's a project instead of a box.
The number is the same. The bottleneck isn't.
显示更多
The "cheap AMD alternative" to the $4,500 NVIDIA DGX Spark costs $3,999 fully specced. That's not an alternative. That's the same price.
Three compact machines are being sold as the answer to local AI: the DGX Spark at ~$4,500, the AMD Strix Halo at ~$4,000, and a Mac at roughly half of either. All three marketed on the same promise — run big models at home, skip the cloud.
But the benchmarks everyone shares measure the wrong thing.
They race token generation — the speed you watch text appear — and on that number all three nearly tie. The Mac generates at 56 tokens/sec. The $4,500 DGX? 84. Barely a gap you'd feel in a chat.
The number that actually separates these machines is prefill — how fast each one reads your context before it answers. And there the gap is brutal: the DGX hits 2,107 tokens/sec, six times the AMD box.
So the real question was never "which is fastest." It's "what's your bottleneck":
- Chatting and short prompts → generation matters, and the cheaper machine already wins
- Long context, RAG, agents, huge files → prefill matters, and only the DGX delivers
- Big models no consumer GPU can hold → AMD's 128GB unified memory, up to 96GB as VRAM
One honest test on the same 30B model settled all three. The full breakdown — prefill vs generation, price vs bottleneck — is pinned 👇
显示更多
"Your Mac is useless for AI without CUDA, buy the $4K DGX." Half right — and the half that's wrong costs you $4,000.
CUDA matters for training and some frameworks. Real. But "running local models"? I benchmarked a Mac Mini against the DGX on the same 30B: generation was 56 vs 84 t/s. Not useless — usable, faster than you read, no CUDA required. llama.cpp and Ollama don't care what logo is on the chip.
Here's the honest split the video skips:
- Commercial fine-tuning, CUDA-only pipelines, big-context prefill → yes, the DGX earns its price
- Running and chatting with local models → your Mac already does this fine
The DGX isn't overpriced. It's mis-pitched. It's a prefill-and-training machine, not a "your Mac is trash" machine. Buy it for the workload that needs it — not out of CUDA FOMO.
Full 3-way benchmark (DGX vs Strix Halo vs Mac Mini, prefill vs generation) pinned.
显示更多
Everyone says the DGX Spark is "2-5x faster" than the $2,000 AMD box. I benchmarked both. It's 6x on prompt processing — and near-identical on generation. Which means most people are about to overpay by $2,700.
Same 30B model, both machines:
Prompt processing (reading your context):
- AMD Strix Halo: 342 t/s
- DGX Spark: 2,107 t/s → 6x
Token generation (writing the answer, the speed you feel):
- AMD Strix Halo: 73 t/s
- DGX Spark: 84 t/s → basically a tie
Read that twice. On the number you actually watch happen — text appearing on screen — a $2,000 box ties a $4,700 one. The Spark's whole premium lives in one place: prefill.
So the buy is simple once you know your bottleneck:
- You chat, you draft, you code with short prompts → Strix Halo. Same feel, pocket $2,700, and it runs Windows and games on the side.
- Your work is huge context — long agent runs, RAG over hundreds of docs, giant files → that 6x prefill is the Spark earning its price. Nothing else touches it.
One tax the AMD hype skips: 128GB is glorious until a tool ships CUDA-only and you're debugging ROCm at midnight. The memory is real. The software lottery is too.
Buy for your bottleneck, not the biggest number on the box.
(Full breakdown of every local AI machine by bottleneck — Spark, Mac, used GPUs, this — pinned.)
显示更多