a $1,999 mini-PC runs a 120-billion AI model at home. ask it to run a 70B and it chokes

the Framework Desktop and its Ryzen AI Max+ 395 twins hand you ~96GB of GPU memory for two grand — enough to load models no single consumer GPU can…

Aliteq
Ravi Malhotra · Hardware Editor

Why a 120B model runs but a 70B model crawls

Token generation speed on these machines is, roughly, memory bandwidth divided by how much of the model has to be read for each token. A dense 70B model reads all 70 billion parameters every token —…

The bandwidth wall — and the prompt-processing catch

That ~212 GB/s figure is the whole story. A discrete RTX 5090 has roughly 1,792 GB/s; a Mac Studio M3 Ultra about 819 GB/s. Strix Halo trades raw bandwidth for sheer capacity, and that trade shows…

So why is it still a bargain? Because 2026 broke every alternative

Normally you would just buy a real GPU. But in mid-2026 the alternatives collapsed. The RTX 5090 has only 32GB and its street price sits above $5,000 — you would need three of them (~$15,000, plus a…

Aliteq

Read the full story

a $1,999 mini-PC runs a 120-billion AI model at home. ask it to run a 70B and it chokes

Read the full story on Aliteq