
Got an RTX 5090? The Best Local LLMs for 32GB of VRAM in 2026
32GB does not unlock 70B models. It buys better quants and a much longer memory for the 27B to 35B class. Five picks that fit a 5090, with the math shown.
Chiptune · 4d ago · 10 min
22 articles · newest first

32GB does not unlock 70B models. It buys better quants and a much longer memory for the 27B to 35B class. Five picks that fit a 5090, with the math shown.
Chiptune · 4d ago · 10 min

The $50 single-server colo deals include a fraction of the power a GPU box needs. What hosting a 1, 2 or 4 GPU machine really costs a month, per kilowatt, from providers' own price pages, and when an office closet or a rented GPU beats it.
0xLax · Sep 27 · 11 min

Should a 10–50 person company buy a GPU server, rent one by the hour, or just pay per token? The three-year bill for each, from live rental prices, today's hardware prices and government power statistics. The answer for most small companies is not the one GPU vendors sell.
0xLax · Sep 27 · 13 min

Scout is mixture-of-experts, so you're buying for 109B in memory, not 17B active. The honest GPU ladder — why a 32GB card is the sweet spot, when 24GB works, and when unified memory beats them all.
Voltage · Sep 22 · 8 min

Renting vs owning a GPU for AI isn't a taste debate — it's a break-even calculation. Here's the honest 2026 math: what an hour in the cloud costs, what a card costs to own, and the crossover where each one wins.
Tensor · Sep 22 · 8 min

Two used RTX 3090s give you 48GB of VRAM for a third of a 5090's price — so is a dual-GPU build the smart move for local AI? The honest math, whether models really split across cards, and the power-and-complexity costs the sticker price hides.
Tensor · Sep 22 · 8 min

At ~$4,600 for 32GB of VRAM, the RTX 5090 is either the only card that makes sense — or a tax on headroom you'll never use. An honest, VRAM-first buy call with the cheaper alternatives priced in.
Tensor · Sep 22 · 8 min

AI demand spiked consumer GPU prices into a mid-2026 peak — and they're cooling off it now, unevenly. The current numbers vs MSRP, why the RTX 5090 won't budge, and what's actually worth buying.
Voltage · Sep 22 · 7 min

I priced out every rung of the local-AI hardware ladder this week. One tier is quietly the best deal on the list, and one isn't worth it at any price.
Voltage · Aug 27 · 9 min

renting looks cheap by the hour. buying looks expensive by the receipt. only one of those numbers is honest
Ledger · Aug 22 · 7 min

somewhere between the memory shortage and the markup, Nvidia's flagship broke pricing so badly that skipping the DIY build is the actual smart move
Voltage · Aug 22 · 7 min

the discount headline is real, but it's not the interesting part. do the math on what the GPU alone costs right now and this deal gets a lot more interesting.
Ledger · Aug 16 · 6 min

The one rule is buy VRAM, not benchmarks — and 2026's price surge is now cooling. My current, honest map: which card for which budget (Arc B580 to RTX 5090), what each actually runs, and when to rent or go unified-memory instead.
Tensor · Aug 14 · 11 min

AMD's own numbers put its workstation GPU at 53 tokens a second on Meta's new 30B model. Nvidia's flagship does more — for over three times the price.
Voltage · Aug 11 · 6 min

For four days in Texas you could buy an RTX 5090 at launch price. This week, Newegg's median RTX 5070 costs $240 more than it did in June.
Ledger · Aug 11 · 6 min

QuakeCon 2026 is, right now, the one place on the planet where Nvidia's flagship still costs what it said on the box.
Ledger · Aug 7 · 6 min

South Korea is the canary. TSMC wafer costs and $20 memory chips just pushed the RTX 5090 to 2.5x its sticker price — here's the actual buy-or-skip math for running AI models at home.
Voltage · Aug 4 · 8 min

That is more than the RTX 5090's $1,999 launch price. We checked the launch figure, today's shelf prices, and what a plain RTX 5080 costs instead.
Voltage · Jul 31 · 6 min

The RTX 5090 is the fastest card money can buy. It's also terrible value for pure gaming. Here's the 4K GPU you should actually buy in 2026, and why the $999 card beats the $1,999 one for most people.
Voltage · Jul 31 · 11 min

To run a 70-billion-parameter model at home you need ~48GB of VRAM. Here's the build that gets you there — and why two cheaper cards beat one expensive one.
Voltage · Jul 30 · 9 min

Measured at 102.7 tokens/sec — double a 3090 — with the only 32GB of VRAM in consumer land. Then the street price doubles the MSRP and the whole verdict flips. Both answers, with the math.
Tensor · Jul 22 · 9 min

Two 3090s give you 48GB, but not double the speed for chat, because llama.cpp's default multi-GPU mode makes the cards take turns. What the benchmarks say, rechecked in October 2026.
Voltage · Jul 21 · 11 min