注册并分享邀请链接,可获得视频播放与邀请奖励。

Antid (@antisadh) “A $249 NVIDIA BOX RUNS LLAMA 3 WITH 8 BILLION PARAMETERS AT 5 TOKENS PER SECOND” — TopicDigg

Antid 的个人资料封面
Antid 的头像
Antid
@antisadh
studying AI daily | helping you save & make money with it
加入 February 2023
81 正在关注    1.5K 粉丝
A $249 NVIDIA BOX RUNS LLAMA 3 WITH 8 BILLION PARAMETERS AT 5 TOKENS PER SECOND AND COSTS $5 A MONTH IN ELECTRICITY WHILE OPENAI CHARGES $200 the video shows the jetson orin nano dev board sitting on a keyboard next to its red sparkfun box, the entire setup smaller than a paperback novel and pulling 7 to 25 watts under load the guy fires up ollama on the board and loads llama 3 with 8 billion parameters into GPU memory, takes about 20 seconds the first time then responds instantly after that on every prompt he runs his benchmark code and clocks 4 to 5 tokens per second on a model that 2 years ago required a $5000 desktop GPU to run at all, now sitting on a $249 board that fits in his palm the chip does 67 trillion AI operations per second, runs entirely offline, has zero rate limits and no API keys, and the only ongoing cost is $5 a month in electricity at 15 watts average openai charges $200 a month for chatgpt pro and tightens rate limits every quarter, this box does 80% of the same job for the price of a coffee and never logs a single conversation anywhere most people are still paying monthly rent to a data center in san francisco while a few are running 8 billion parameter models on a board the size of a deck of cards in their bedroom the window is open, follow and bookmark before it closes
显示更多