ALITEQ.

apple's mac studio just beat nvidia's own local-AI chip and the win doesn't add up

Tom's Hardware ran the numbers on three very different local-AI machines, and the Mac won on raw speed almost across the board — just not by the margin its spec sheet promised.

Ravi MalhotraUpdated 2h ago6 min readWeb story
Apple's Mac Studio desktop computer on a desk

Tom's Hardware just ran a head-to-head nobody else has bothered to run properly: the Mac Studio's M4 Max chip against Nvidia's GB10 — the Grace Blackwell chip inside DGX Spark-class mini PCs — and AMD's Strix Halo, all running the same local LLMs. The Mac won on raw decode speed in every single test. That part isn't shocking if you've followed Apple Silicon's local-AI reputation. What's actually interesting is the size of the win: the M4 Max has exactly double the memory bandwidth of both competitors on paper, 546 GB/s versus 273 GB/s each, but that 2x spec advantage didn't turn into a clean 2x real-world win. On one model, it barely beat 25%.

The bandwidth math, side by side

Memory bandwidth: what the spec sheet promises

Mac Studio (M4 Max, 40-core GPU)

Memory bandwidth
546 GB/s
Chip type
Apple Silicon unified memory

Nvidia GB10 (DGX Spark-class)

Memory bandwidth
273 GB/s
Chip type
Grace Blackwell superchip

AMD Strix Halo

Memory bandwidth
273 GB/s
Chip type
Ryzen AI Max unified memory

Where the Mac's win was massive

On Gemma 4 12B, the M4 Max wasn't just faster — it was in a different class. Tom's Hardware measured a 1.8x throughput advantage over the GB10 system and a 2.26x advantage over Strix Halo. That's a bigger gap than the raw bandwidth numbers alone would predict, which cuts against the idea that bandwidth is the only thing that matters here — something about how Gemma 4 12B's architecture maps onto Apple's GPU cores plays to the M4 Max's specific strengths.

Gemma 4 12B: throughput relative to Strix Halo

AMD Strix Halo1x (baseline)

Ryzen AI Max unified memory

Nvidia GB101.26x

Grace Blackwell superchip

Mac Studio M4 Max2.26x

40-core GPU, Apple Silicon

The part that doesn't fully add up

Here's the wrinkle: on Qwen 3.6-35B-A3B, that same M4 Max — same 2x bandwidth advantage on paper — only pulled 25% ahead of the other two. Compare that to the 80-126% lead it built on Gemma 4 12B and something is clearly not just about memory bandwidth. Qwen 3.6-35B-A3B is a mixture-of-experts model; Gemma 4 12B is dense. MoE architectures activate only a fraction of their parameters per token, which changes the compute-to-memory-access ratio in ways that can favor different hardware differently. My take: if you're shopping by a single spec number like memory bandwidth, this benchmark should worry you a little — the same chip's real-world lead swung from 25% to 226% depending purely on which model you ran, the same trap we flagged in our AMD vs Nvidia for local AI breakdown. That's the difference between 'decent upgrade' and 'no contest' for a buyer trying to plan around one number.

Apple Mac Studio desktop computer connected to a monitor
The Mac Studio's M4 Max chip posted the fastest decode throughput of the three systems tested — by a wildly inconsistent margin depending on the model. · Unsplash

What this actually means if you're buying

If you're choosing hardware to run models locally, the lesson isn't 'buy the Mac' or 'buy the PC with the GB10 chip' — it's that a single spec-sheet number won't tell you how your specific model will actually perform. We've made this case before with how much VRAM you actually need for AI models: the honest answer is always 'it depends on the model you're running,' and this benchmark is a clean demonstration of the same principle applied to bandwidth instead of capacity. If you're specifically weighing a MacBook against other hardware for local AI, check the model family you actually plan to run before trusting any single spec.

Mac Studio vs GB10 vs Strix Halo, answered

Is the Mac Studio just better for local AI overall?
On this specific decode-throughput benchmark, yes — it won every test. But 'better' depends on your model, your budget, and whether you need CUDA-specific tooling that only the Nvidia side supports well.
What is Nvidia's GB10 chip, exactly?
GB10 is Nvidia's Grace Blackwell superchip used in DGX Spark-class compact AI systems, like Dell's Pro Max GB10 — a small-form-factor machine aimed at running larger models locally without a full workstation GPU.
What is AMD Strix Halo?
Strix Halo is AMD's Ryzen AI Max platform, which pairs CPU and GPU cores with unified memory in a laptop- or mini-PC-friendly package, competing directly with Apple Silicon and Nvidia's GB10 on the same local-AI-appliance turf.
Does more memory bandwidth always mean faster local AI?
No — this benchmark is the proof. The M4 Max has exactly double the bandwidth of both rivals but only converted that into a 25% real-world lead on one model and a 226% lead on another, depending on the model's architecture.

The honest advice: don't shop for local-AI hardware off a single bandwidth or VRAM number the way you'd shop for a gaming GPU off an FPS chart. Pick the model family you actually plan to run first, then look for benchmarks — like this one — that test that specific architecture, not just a synthetic memory-bandwidth ceiling. We'll keep tracking which of these platforms holds up as more local models ship through the rest of 2026, and if Nvidia or AMD publish their own numbers disputing this, that's worth a follow-up on its own.

Hardware Editor

Ravi Malhotra

Ravi has been building and taking apart PCs since the single-core days — his idea of a good weekend is a repaste and a spreadsheet full of thermals. He covers GPUs, CPUs and the build decisions that actually move frame rates, and he'd rather hand you a benchmark than a press release.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading