The 4090 is 34% faster for local AI. Your electricity meter can't tell them apart

Measured llama.cpp numbers say the 4090 wins every speed test and costs twice as much — while energy per generated token is a dead heat. Here's who…

Aliteq
Ravi Malhotra · Hardware Editor

The short answer

Both cards have 24GB of VRAM, so they run exactly the same models — the ceiling is identical. The RTX 4090 generates tokens ~34% faster (69.8 vs 52.1 t/s on Qwen3-14B Q4_K at 16k, measured in…

Buy the used 3090 if your goal is the cheapest 24GB that runs 30B-class models. It's the value pick, and it isn't close: ~$19 per token/sec vs ~$28 on the 4090.

Buy the 4090 if you feed models long documents all day — prompt processing is 3,452 vs 1,679 t/s at 16k, a 2x gap you feel on every big paste — or if you want native FP8 for vLLM serving.

Energy per generated token is nearly identical (~6.4 vs ~6.7 J/token derived from rated power and measured speed), so "efficiency" shouldn't decide this.

Neither card runs a dense 70B properly. That's a different decision entirely.

Buy the 3090 for value. Buy the 4090 for prefill.

Same 24GB, same model ceiling. The 4090's 34% generation lead and 2x prompt-processing lead are real, measured, and cost roughly double. Energy per token is a wash. Most home users should take the…

Aliteq

Read the full story

The 4090 is 34% faster for local AI. Your electricity meter can't tell them apart

Read the full story on Aliteq