A $249 NVIDIA BOX RUNS LLAMA 3 WITH 8 BILLION PARAMETERS AT 5 TOKENS PER SECOND AND COSTS $5 A MONTH IN ELECTRICITY WHILE OPENAI CHARGES $200
the video shows the jetson orin nano dev board sitting on a keyboard next to its red sparkfun box, the entire setup smaller than a paperback novel and pulling 7 to 25 watts under load
the guy fires up ollama on the board and loads llama 3 with 8 billion parameters into GPU memory, takes about 20 seconds the first time then responds instantly after that on every prompt
he runs his benchmark code and clocks 4 to 5 tokens per second on a model that 2 years ago required a $5000 desktop GPU to run at all, now sitting on a $249 board that fits in his palm
the chip does 67 trillion AI operations per second, runs entirely offline, has zero rate limits and no API keys, and the only ongoing cost is $5 a month in electricity at 15 watts average
openai charges $200 a month for chatgpt pro and tightens rate limits every quarter, this box does 80% of the same job for the price of a coffee and never logs a single conversation anywhere
most people are still paying monthly rent to a data center in san francisco while a few are running 8 billion parameter models on a board the size of a deck of cards in their bedroom
the window is open, follow and bookmark before it closes
显示更多