ALITEQ.

Everyone's buying $4,700 'AI desktops' right now. A lab test just showed most people need way less

StorageReview ran real benchmarks on every local-AI desktop you can buy in 2026 — and the winner isn't what most shoppers assume.

Ravi MalhotraUpdated 1h ago7 min readWeb story
A compact desktop PC built for running local AI models at home

The first independent lab test of 2026's 'AI desktop' category is out, and it settles an argument that's been running in local-AI forums for months: do you need a tower stuffed with discrete GPUs, or does a small unified-memory box actually get the job done? StorageReview's lab-tested leaderboard, published August 13, ran real vLLM serving benchmarks — time-to-first-token, tokens-per-second, sustained throughput — across seven systems, from a $4,699 shoebox-sized AI computer to an eight-GPU liquid-cooled beast. The winner isn't the machine most shoppers are about to impulse-buy.

The lab didn't rank spec sheets

StorageReview's methodology is the useful part here: systems were scored on a composite of vLLM online serving throughput, time-to-first-token, and time-per-output-token across GPT-OSS-120B, Llama 3.1 8B, Mistral Small 3.1 24B, and Qwen3 Coder 30B, plus MAMF compute efficiency and storage throughput under GDSIO. That's a meaningfully different exercise than reading a spec sheet and multiplying memory bandwidth by core count. A machine that looks strong on paper can still choke on real serving latency, and vice versa.

Seven systems, one real fork in the road

Every system on the leaderboard falls into one of two camps. Unified-memory systems — the DGX Spark, Acer's Veriton GN100, AMD's Ryzen AI Max boxes — pool CPU and GPU memory into one big pot, up to 128GB, at appliance prices. Discrete-VRAM towers — Dell's Precision 7875, HP's Z8 Fury G6i, the Comino Grando — stack multiple RTX PRO 6000 cards for far more raw memory and throughput, at far higher cost. The dividing line is model size: a 70B model at 4-bit quantization needs roughly 40-48GB of usable memory before you add any context window, which every unified-memory system here clears comfortably. Push past 200B parameters and you're firmly in multi-GPU tower territory.

What actually made the leaderboard

NVIDIA DGX Spark

System
128GB unified LPDDR5X
Memory
$4,699
Price
Best overall deskside AI system

Acer Veriton GN100

System
128GB unified LPDDR5X
Memory
Not widely listed yet
Price
Best GB10 implementation (thermals)

AMD Ryzen AI Max 'Strix Halo'

System
Up to 128GB unified LPDDR5X
Memory
Varies by OEM
Price
Best x86 alternative, handles 200B-class models

HP Z2 Mini G1a

System
128GB unified (up to 96GB as VRAM)
Memory
~$3,300–3,400
Price
Best without a discrete GPU

Dell Precision 7875

System
192GB GDDR7 (dual RTX PRO 6000)
Memory
Workstation pricing
Price
Fastest standard tower

HP Z8 Fury G6i

System
Up to 384GB (4× RTX PRO 6000)
Memory
Workstation pricing
Price
Best multi-GPU platform

Comino Grando

System
768GB GDDR7 (8× RTX PRO 6000)
Memory
Enterprise pricing
Price
The extreme pick
A compact AI desktop PC on a desk
Unified-memory mini PCs like the DGX Spark and Z2 Mini G1a are reshaping what a 'local AI desktop' even means. · Unsplash

Where the multi-GPU towers actually pay off

Dell's Precision 7875, HP's Z8 Fury G6i, and the Comino Grando aren't for the person asking 'can I run a chatbot on my desk.' They're for 200B-plus parameter models, multiple concurrent users, or production serving — and they cost accordingly. Given how DRAM and GDDR7 shortages have driven GPU prices up through 2026, a tower built on multiple RTX PRO 6000 cards is a worse value proposition today than it was in January, purely on memory-chip cost. Before spending that kind of money, it's worth actually pricing out whether a single RTX PRO 6000 is worth $16,000 for your workload in the first place, let alone four of them.

Verdict

The buy call

Most home local-AI users should look at the Ryzen AI Max / unified-memory tier before a multi-GPU tower. It clears the 40-48GB threshold that covers the vast majority of useful open models, costs a fraction of a multi-GPU workstation, and — per this leaderboard — genuinely performs.

Best for: Home users and solo developers → unified memory. Teams serving 200B+ models to multiple people → discrete-VRAM towers.

Do I need a discrete GPU to run local AI in 2026?
No — StorageReview's testing shows unified-memory systems like the HP Z2 Mini G1a can run GPT-OSS 120B with zero discrete GPU, though a dedicated card still wins on raw throughput for smaller models.
How much memory do I actually need?
Roughly 40-48GB of usable memory for a 70B-parameter model at 4-bit quantization, before you add context. Below that, you're limited to 20-30B-class models.
Is the $4,699 DGX Spark worth it?
Only if you specifically need Nvidia's CUDA stack and GB10-specific tooling. Cheaper Ryzen AI Max systems land within striking distance on real-world throughput for most model sizes.
What's the cheapest real entry point into local AI?
A single midrange GPU with 16GB of VRAM, covered in our best GPU for local AI breakdown, still beats any system on this leaderboard on dollars-per-token for models under 20B parameters.

None of this is a call to buy the biggest box you can afford. It's a call to actually match memory and price to the model size you'll realistically run — something the DIY-versus-prebuilt local AI PC math has been arguing since before this leaderboard existed. StorageReview just put real numbers behind it.

Hardware Editor

Ravi Malhotra

Ravi has been building and taking apart PCs since the single-core days — his idea of a good weekend is a repaste and a spreadsheet full of thermals. He covers GPUs, CPUs and the build decisions that actually move frame rates, and he'd rather hand you a benchmark than a press release.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading