NVIDIA packed 128GB of unified memory into a 1.13L box called DGX Spark.
Kevin tore one apart today.
Inside: a GB10 Superchip. 20 ARM cores. NVLink fusing CPU and GPU into 1 shared pool. Models that need a rack now fit in something smaller than a wine bottle.
Then the weak spot.
NVIDIA shipped it with a 2242 SSD. Tiny stick. Poor performance. Not rated for these workloads. A 2280 slot would have cost them nothing.
But look closer at the board.
2 ports. 200 gig each. The NIC shrunk to a fraction of a normal card.
That changes the math.
Skip the bad SSD. Run NVMe over fabrics to a real data center early tests already pulled 100 gig, with drivers still raw. Or cable 2 Sparks straight into each other: 256GB of unified memory across both, shard the model with vLLM, done.
Now scale it. Hang a stack of them on a 200-gig switch. Full bandwidth per unit. Data center storage behind them. The weak drive stops mattering.
People used to rack trays of Mac Minis.
The next tray holds a data center.
显示更多