
a 5-year-old GPU still beats nvidia's newest budget card for local AI — here's the catch
24GB of old VRAM versus 16GB of new, low-power silicon — I ran the real math on what each one costs, up front and over a year.
Ravi Malhotra · 6d ago · 6 min
11 articles · newest first

24GB of old VRAM versus 16GB of new, low-power silicon — I ran the real math on what each one costs, up front and over a year.
Ravi Malhotra · 6d ago · 6 min

Four real paths to running a 70B-parameter model on your own hardware or by the hour — priced out with actual 2026 numbers.
Lena Fischer · 6d ago · 8 min

A third price hike in one year is reportedly coming for the RTX 50 series — and if you're buying VRAM for local AI, not frame rates, the calculation isn't the same as 'wait for a sale.'
Ravi Malhotra · 6d ago · 7 min

For local AI, VRAM beats speed every time. A slower card with more memory runs bigger models than a faster one that can't fit them. Here's the best GPU for local AI at every budget.
Ravi Malhotra · Aug 1 · 10 min

A used 3090 is still the best VRAM-per-dollar card for local AI — but many were mining cards, and the memory runs hot. Here's exactly what to check before you hand over the cash.
Ravi Malhotra · Jul 28 · 9 min

Qwen3-32B at Q4 needs about 20GB of VRAM, which rules out every 16GB card and makes this a simple question: what's the cheapest 24GB GPU that runs it well? The answer is used.
Ravi Malhotra · Jul 28 · 9 min

Llama 70B needs about 40GB of VRAM to run well — and almost no single consumer card has it. Here's what actually works, visualized, with the measured speeds and real costs for each path.
Ravi Malhotra · Jul 23 · 10 min

A new RTX 5060 Ti 16GB looks like the safe budget pick. But 16GB can't load the 32B models a used 3090 runs at 30 tokens/sec — and in local AI, the card that opens the file wins before speed matters.
Ravi Malhotra · Jul 22 · 9 min

On small-model generation they're within 5% of each other — measured. The real decision lives in what each machine physically can't do, and nobody selling either will tell you.
Lena Fischer · Jul 22 · 10 min

Measured llama.cpp numbers say the 4090 wins every speed test and costs twice as much — while energy per generated token is a dead heat. Here's who each card is actually for.
Ravi Malhotra · Jul 22 · 9 min

The dual-3090 build doesn't give you double the bandwidth for chat — and the reason has nothing to do with the cards. What the measurements actually say.
Ravi Malhotra · Jul 21 · 11 min