JEFF GEERLING RAN DEEPSEEK 671B AT HOME AND THE SETUP KILLS A $400/MONTH OPENAI AND CLAUDE CODE STACK FOR $9 A MONTH
jeff geerling is an independent homelab developer who runs his entire AI workload from a server in his basement
he wired a 192 core ampere one server to an AMD W7700 GPU and pointed ollama at deepseek 671B, 400 watts at the wall, 50 tokens per second
a heavy developer pays $400 a month across claude code max and chatgpt pro, jeff's setup runs the same workloads locally for $9 a month in electricity
every benchmark, every power draw number and every config file is on his blog for free, the math is public, the build is reproducible, the cost is real
the week he posted the rack nvidia lost $500 billion in market cap, the moat openai sold for $200 a month was always one local box away from being gone
the window is open, follow and bookmark before it closes
显示更多
AMD ENGINEER'S PALM-SIZED MINI PC RUNS 235B MODELS FOR $9/MONTH AND KILLS A $200/MONTH CLAUDE CODE AND OPENAI SUB
at CES 2026 in las vegas, AMD CEO lisa su walked on stage with a small black box behind her, not a server, not a data center render, a mini PC the size of a hardcover book
a few months later in shanghai she walked up to that same device and signed it with her name, the box is the gmktec EVO-X2, $1,700 once, AMD ryzen AI max+ 395 inside
the chip is the first x86 silicon ever built that runs a 200 billion parameter model on one piece of hardware, 128GB unified memory, 110GB usable VRAM on linux, no separate graphics card
it runs qwen 3 235B fully and smoothly, kills a $440/month stack of claude code, chatgpt, gemini and cursor for $9 a month in electricity, no per token fees, nothing ever leaves the machine
the window is open, follow and bookmark before it closes
显示更多
AN AMD ENGINEER SHIPS A PALM-SIZED MINI PC THAT RUNS 235B MODELS FOR $9/MONTH AND KILLS A $200/MONTH OPENAI OR CLAUDE CODE SUB
at CES 2026 in las vegas, AMD CEO lisa su walked on stage with a small black box behind her, not a server, not a data center render, a mini PC the size of a hardcover book
a few months later in shanghai she walked up to that same device and signed it with her name, the box is the gmktec EVO-X2, $1,700 once, AMD ryzen AI max+ 395 inside
the chip is the first x86 silicon ever built that runs a 200 billion parameter model on one piece of hardware, 128GB unified memory, 110GB usable VRAM on linux, no separate graphics card
it runs qwen 3 235B fully and smoothly, plus deepseek v3 and llama 3.3 70B with no quantization, kills a $440/month claude code, chatgpt, gemini and cursor stack for $9 in electricity
setup takes 3 commands and 15 minutes, ollama loads the model, claude code points to localhost with one environment variable, same interface, zero per token fees, nothing ever leaves the machine
the window is open, follow and bookmark before it closes
显示更多
KEVIN BUILT A PRIVATE AI SERVER FROM 5 MAC MINIS PULLING UNDER 30 WATTS IDLE AND KILLED HIS $200/MONTH CLAUDE BILL FOR $3 A MONTH IN ELECTRICITY
two months ago kevin posted his claude code bill on reddit, $170 in 10 days of building a saas, the quality was magic but the bill was not
the top reply said "i bought a mac mini m4, haven't paid anthropic since" and the thread exploded with developers doing the same math
uber rolled out claude code to 5,000 engineers and watched bills hit $500 to $2,000 per person, burning through their entire $3.4 billion AI budget in 4 months
the mac mini m4 starts at $599 one time, pulls 10 to 20 watts running 24/7, and costs $3 a month in electricity while a windows AI machine doing the same job costs $30 to $50 a month just to stay on
ollama added support for the anthropic messages api in january 2026, so claude code itself connects to your local mac mini with one environment variable, same interface, zero API costs
apple stores ran out of mac minis in 2026 because $599 one time beats $200 a month forever, that shortage is the most honest product review any machine has ever received
the window is open, follow and bookmark before it closes.
显示更多
JENSEN HUANG BRAGGED NVIDIA HAS THE LOWEST COST PER TOKEN IN THE WORLD. DEVELOPERS DROPPED THEIR COST TO $0 WITH A $599 MAC MINI
two months ago a developer posted his claude code bill on reddit, $170 in 10 days of building a saas, the quality was magic but the bill was not
the top reply said "i bought a mac mini m4, haven't paid anthropic since" and the thread exploded with developers doing the same math
uber rolled out claude code to 5,000 engineers and watched bills hit $500 to $2,000 per person, burning through their entire $3.4 billion AI budget in 4 months
ollama added support for the anthropic messages api in january 2026, so claude code itself connects to your local mac mini with one environment variable, same interface, zero API costs
apple stores ran out of mac minis in 2026 because $599 one time beats $200 a month forever, that shortage is the most honest product review any machine has ever received
the window is open, follow and bookmark before it closes
显示更多
A $249 NVIDIA BOX RUNS LLAMA 3 WITH 8 BILLION PARAMETERS AT 5 TOKENS PER SECOND AND COSTS $5 A MONTH IN ELECTRICITY WHILE OPENAI CHARGES $200
the video shows the jetson orin nano dev board sitting on a keyboard next to its red sparkfun box, the entire setup smaller than a paperback novel and pulling 7 to 25 watts under load
the guy fires up ollama on the board and loads llama 3 with 8 billion parameters into GPU memory, takes about 20 seconds the first time then responds instantly after that on every prompt
he runs his benchmark code and clocks 4 to 5 tokens per second on a model that 2 years ago required a $5000 desktop GPU to run at all, now sitting on a $249 board that fits in his palm
the chip does 67 trillion AI operations per second, runs entirely offline, has zero rate limits and no API keys, and the only ongoing cost is $5 a month in electricity at 15 watts average
openai charges $200 a month for chatgpt pro and tightens rate limits every quarter, this box does 80% of the same job for the price of a coffee and never logs a single conversation anywhere
most people are still paying monthly rent to a data center in san francisco while a few are running 8 billion parameter models on a board the size of a deck of cards in their bedroom
the window is open, follow and bookmark before it closes
显示更多
ONE WEEKEND AND $249 IN HARDWARE TURNED A 3 GIGABYTE AI MODEL INTO AN OFFLINE CHATGPT THAT COSTS $5 A MONTH WHILE OPENAI CHARGES $200
the video shows a small lilliput monitor sitting on a sparkfun box with the jetson orin nano dev board wired up underneath, the whole setup costs less than a single month of chatgpt pro
the guy spent saturday installing the OS and using docker images to run microsoft's phi-2 model locally, then he built a simple script that lets him chat with it directly on the device
he asks the box for a joke and the response comes back instantly with full chat history saved, so he can ask for another one and it remembers the context just like chatgpt does on the cloud
the model is only 3 gigabytes and runs at conversational speed on a chip that fits in your palm and pulls 7 to 25 watts, while a regular PC running AI work pulls 300 to 500
openai charges $200 a month for chatgpt pro and rate limits you at the worst moments, this box does 80% of the same job for $5 a month in electricity with no logging, no API key, no cloud server
most people are still renting AI from someone else's data center while a few are running it on a $249 board in their bedroom
the window is open, follow and bookmark before it closes.
显示更多