Measured llama.cpp numbers say the 4090 wins every speed test and costs twice as much — while energy per generated token is a dead heat. Here's who each card is actually for.
Here's the part that surprised me when I actually sat down with the measurements. Hardware-corner.net's llama.cpp benchmarks — same harness, same quant, same 16k context — put the RTX 4090 at 69.8 tokens/sec generating with Qwen3-14B Q4_K, against 52.1 for the RTX 3090. That's the 34%. Now do the electricity: the 4090 is rated at 450W, the 3090 at 350W. Divide power by speed and you get roughly 6.4 joules per token on the 4090 and 6.7 on the 3090. Per token generated, five years of architectural progress bought you almost nothing. The 4090 finishes sooner; your meter barely notices the difference.
That reframes this whole comparison, because it kills the vague "the 4090 is more efficient" argument you'll read everywhere — and leaves the real question standing alone: is 34% more speed worth roughly 2x the price? For most people running models at home, I don't think it is. But there are two buyer types where the 4090 is clearly right, and I'll get to both.
Same ceiling, different speed — why VRAM makes this comparison weird
GPU-vs-GPU pages usually argue about which card is "more capable." Here that's settled by a spec sheet: NVIDIA lists both at 24GB — GDDR6X on both, 936 GB/s of bandwidth on the 3090, 1,008 GB/s on the 4090. Every model that fits one fits the other. Qwen3-32B at Q4_K? Fits both, tightly. Gemma 3 27B? Both. A dense 70B at Q4? Neither, not really — that's roughly 40GB with context, and you're into offload tricks or multi-GPU territory on either card.
So capability is a tie by construction. What you're actually buying is speed, and the interesting part is that the speed gap is bigger than the bandwidth gap. 1,008 vs 936 GB/s is 8%. The measured generation gap is 34%. Token generation is supposed to be bandwidth-bound, so where do the other 26 points come from? Mostly the 4090's enormous L2 cache — 72MB against the 3090's 6MB — which means a meaningful slice of weight reads never touch the memory bus at all. It's the one Ada advantage that shows up every single token.
What the vendor sheets and the measurements say
VRAM
RTX 3090
24GB GDDR6X
RTX 4090
24GB GDDR6X
Memory bandwidth
RTX 3090
936 GB/s
RTX 4090
1,008 GB/s
L2 cache
RTX 3090
6MB
RTX 4090
72MB
Rated power
RTX 3090
350W
RTX 4090
450W
Gen speed, Qwen3-14B Q4_K @16k*
RTX 3090
52.1 t/s
RTX 4090
69.8 t/s
Prompt processing @16k*
RTX 3090
1,679 t/s
RTX 4090
3,452 t/s
Gen speed, Qwen3-32B Q4_K @16k*
RTX 3090
30.3 t/s
RTX 4090
—
Tracked used price (US)*
RTX 3090
~$1,000
RTX 4090
~$2,200
Price per token/sec (14B)*
RTX 3090
$19.18
RTX 4090
$27.95
Native FP8 (Tensor Cores)
RTX 3090
No
RTX 4090
Yes
RTX 3090
RTX 4090
VRAM
24GB GDDR6X
24GB GDDR6X
Memory bandwidth
936 GB/s
1,008 GB/s
L2 cache
6MB
72MB
Rated power
350W
450W
Gen speed, Qwen3-14B Q4_K @16k*
52.1 t/s
69.8 t/s
Prompt processing @16k*
1,679 t/s
3,452 t/s
Gen speed, Qwen3-32B Q4_K @16k*
30.3 t/s
—
Tracked used price (US)*
~$1,000
~$2,200
Price per token/sec (14B)*
$19.18
$27.95
Native FP8 (Tensor Cores)
No
Yes
*Measured figures and tracked prices from hardware-corner.net's RTX 3090 and RTX 4090 LLM benchmark pages — their llama.cpp benchmark suite (the 3090 page lists CUDA with flash attention enabled; benchmarks dated March 2026). We don't run our own benchmarks; theirs are the best harness-disclosed numbers publicly available for these cards. Used prices move around a lot right now — more on that below.
The prompt-processing gap is the one people underrate
Everyone benchmarks generation speed because it's the number you watch while the model types. But if your workflow is "paste a long document, ask questions about it" — and for coding assistants and RAG that's basically the whole job — the number that sets your wait is prompt processing. There the 4090 doubles the 3090: 3,452 vs 1,679 tokens/sec at 16k context in the same measurements.
In practice: handing a model a 16,000-token codebase chunk takes the 3090 about ten seconds before the first word of the reply. The 4090 does it in under five. Do that forty times a day and the 4090 stops feeling like a luxury. This is honestly the strongest pro-4090 argument, and it's the one the spec-sheet comparisons mostly skip because it doesn't show up in a bandwidth column.
Measured llama.cpp throughput — Qwen3-14B Q4_K at 16k context (hardware-corner.net, March 2026)
RTX 3090 — generation52.1 t/s
RTX 4090 — generation69.8 t/s
RTX 3090 — prompt processing1,679 t/s
RTX 4090 — prompt processing3,452 t/s
The 2026 price problem — and why every static comparison is wrong
In early 2025 you could find used 3090s at $700 and this comparison wrote itself. Then the memory shortage happened to everything with DRAM in it, and used cards followed new ones up. Hardware-corner's tracker currently averages the 3090 at ~$1,000 and the 4090 at ~$2,200 on the US used market. Any page quoting $650 3090s is describing a market that no longer exists.
On used-3090 risk, since I'd be annoyed if a guide didn't say it: a lot of these cards mined for a living, and the GDDR6X on the 3090 runs famously hot on the back of the board. Ask for HWiNFO screenshots under load before buying; memory-junction temperatures sustained above ~100°C mean the thermal pads are overdue. Budget $30 and an hour for a repad on any card without service history. It's not a reason to avoid the 3090 — it's a reason to buy from someone who can answer questions.
Where the 4090 genuinely earns its premium
You process long prompts constantly. Coding assistants, document Q&A, RAG pipelines. The 2x prefill gap is your daily experience, not a benchmark row.
You want FP8. Ada's Tensor Cores do native FP8; Ampere's don't. For llama.cpp chat this barely matters today, but if you're serving models with vLLM or care about the formats labs are actually shipping, it's a real fork in the road.
You also generate images or video. Diffusion is compute-bound, not bandwidth-bound — the 4090's advantage there is far larger than 34%, and the 3090 falls well behind.
Resale. The 4090 will still be a current-ish card in 2028. The 3090 will be eight years old.
And one honest mark against the 4090: at ~$2,200 used it now overlaps with what two 3090s plus a beefier PSU cost — and 48GB changes what you can run, not just how fast. If your ambitions include 70B-class models, that comparison matters more than this one.
Six years old, still the cheapest honest 24GB in the game. · NVIDIA
My call
I keep coming back to the fact that both cards run the same models. When capability is tied, price-per-performance decides, and $19 versus $28 per token/sec isn't subtle. The 3090 remains the card I'd point most people at for a first serious local-AI box — check what actually fits in 24GB and you'll find the ceiling is generous: Qwen3-32B at ~30 tokens/sec is a perfectly usable daily model.
Verdict
Buy the 3090 for value. Buy the 4090 for prefill.
Same 24GB, same model ceiling. The 4090's 34% generation lead and 2x prompt-processing lead are real, measured, and cost roughly double. Energy per token is a wash. Most home users should take the ~$1,000 used 3090 and put the difference toward RAM, a better PSU — or 6,800 hours of rented GPU time.
Best for: 3090: first local-AI build, chat/agent workloads, value hunters. 4090: long-context daily drivers, vLLM/FP8 servers, anyone also doing image gen.
The questions people actually ask
Can the RTX 4090 run bigger models than the 3090?
No. Both are 24GB cards and the model ceiling is identical — Qwen3-32B or Gemma 3 27B at Q4 is the practical top on either. The 4090 runs the same models faster; it doesn't run larger ones. If you need bigger, the answer is two cards, an offload strategy, or renting — not a 4090.
Is a used RTX 3090 risky in 2026?
Manageable, not trivial. Many were mining cards, and 3090 GDDR6X memory runs hot enough that aged thermal pads are the most common real problem. Ask for temperature screenshots under load, prefer sellers with service history, and budget for a $30 thermal-pad replacement if there's none. The GPU silicon itself rarely dies from mining; cooling neglect is what kills these cards.
Should I wait for a 24GB RTX 50 Super instead?
NVIDIA reportedly finished a 24GB 50-series Super and then shelved it over GDDR7 supply costs — we covered the reporting separately. Waiting on an unannounced card in a memory shortage is a bet with no date attached. If you need a GPU this year, buy for this year's market.
What about a 3090 Ti?
Genuinely interesting middle ground when priced near the 3090: same 24GB but 1,008 GB/s of bandwidth — the 4090's number — and better-behaved memory cooling. They're rarer used, and don't pay 4090 money for one: it still lacks FP8 and the giant L2, so it splits the generation gap rather than closing it.
If you're still deciding what you'd actually run on 24GB, our VRAM calculator does the fit math per model and quant, and the cost-to-run pages keep live rent-vs-buy numbers so you can check whether owning beats renting for your hours. Both get you to the same place: know your workload first, then buy the card it needs — not the card the benchmarks flatter.