NVIDIA spent three months optimising the DGX Spark. Token generation got slower.

Same machine, same benchmarks, three months apart. Prefill improved 27%. Generation regressed on every model — and that tells you exactly who should…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short answer

The DGX Spark is a strong prefill machine and a mediocre token generation machine, and almost everyone evaluating it is looking at the wrong number. Its 128 GB of unified memory runs on a 273 GB/s…

Buy it for long-context, agentic, MoE-heavy work, where it beats an M2 Ultra at prompt processing by 1.48× rising to 2.17× at 32k context.

Don't buy it for dense-model chat. On a dense 7B it manages 29.43 t/s against an M2 Ultra's 79.68 — almost exactly the bandwidth ratio.

NVIDIA advertises *1 PFLOP FP4* and *models up to 200 billion parameters*. Llama-3.1-70B was measured at 2.7 tokens/sec decode. It runs. You wouldn't use it.

No software update will change the decode number. Three months of optimisation already tried.

Right machine, oversold on the wrong number

The DGX Spark is a legitimately good long-context MoE and prototyping box with a real CUDA stack in a small quiet package. It is a poor dense-model chat machine and nothing will change that. The…

Aliteq

Read the full story

NVIDIA spent three months optimising the DGX Spark. Token generation got slower.

Read the full story on Aliteq