ALITEQ.

how to turn 128GB of RAM into 96GB of AI memory: the Strix Halo VRAM trick nobody explains

AMD's Ryzen AI Max+ 395 lets you hand up to 96GB of its unified memory to the GPU — enough to run a 70-billion-parameter model on a mini-PC. Here's exactly how much you can allocate, how to do it, and the catch.

Ravi MalhotraUpdated 2h ago9 min read
Macro shot of a system-on-chip on a circuit board, representing the Strix Halo APU

The question people keep asking about AMD's Strix Halo chip is deceptively simple: it has up to 128GB of memory, but how much of that can I actually use as VRAM for AI? The answer is up to 96GB — and that single number is why a $1,500–2,000 mini-PC can run a 70-billion-parameter model that a $2,000 graphics card physically cannot. But the way you get there isn't obvious, the allocation is set somewhere most people never look, and there's a bandwidth catch that decides whether this is the right machine for you. Let me walk through all three.

How much you can actually allocate

On a 128GB Strix Halo system, AMD's Variable Graphics Memory lets you convert up to 96GB into VRAM, leaving roughly 32GB for the operating system and everything else. That's the ceiling; you can also choose less. The reason it's not the full 128GB is that the CPU and OS still need working memory — hand the GPU everything and the system falls over. 96GB is the sweet spot AMD and the mini-PC vendors settled on, and it's a colossal amount of AI memory by consumer standards: four times a 5090, and available in a box the size of a paperback.

The mechanism matters. Unlike a discrete GPU with its own dedicated VRAM chips, Strix Halo uses a Unified Memory Architecture — the CPU, the RDNA 3.5 integrated GPU, and system RAM all draw from the same physical pool. So 'allocating VRAM' really means deciding how much of your shared memory the GPU gets to claim. It's the same idea that makes Apple's Mac Studio able to run huge models: when memory is unified, capacity is a setting, not a soldered-on limit.

What 96GB unlocks — models a 24GB card can't hold
on 96GB Strix Halo needs ≈58 GB
RTX 4090 (24GB)24 GBover 34 GB
RTX 5090 (32GB)32 GBover 26 GB
Strix Halo (96GB)96 GBfits
Sized against Llama 70B at Q6 (~58GB). A 24GB discrete GPU can't hold it; Strix Halo's 96GB clears it with room for context. That capacity is the entire pitch.

How to set it (the step people miss)

Here's what trips people up: the VRAM split is configured in the BIOS/UEFI at boot, not on the fly in Windows. You reboot into firmware settings, find the Variable Graphics Memory (sometimes labelled UMA Frame Buffer Size or a similar name depending on the vendor), and set how much memory to dedicate to the GPU. Some systems also expose a split via AMD's Adrenalin software, but the firmware setting is the authoritative one. Set it, save, reboot — and your allocation shows up as GPU memory. If you allocate 96GB and then wonder why Windows says you only have 32GB of RAM, that's working as intended: you gave the other 96 to the GPU.

A compact desktop computing setup on a white desk
A Strix Halo mini-PC puts 96GB of AI memory in a box you can hold — the appeal is capacity in a tiny, quiet, efficient package. · Unsplash

The catch: bandwidth, not capacity

Now the honest part. Having 96GB of VRAM lets you load enormous models; how fast they run is a different question, and it's where Strix Halo trades away its advantage. Its unified memory runs at roughly 256 GB/s. A discrete RTX 4090 does over 1,000 GB/s; a 5090 nearly 1,800. Since token generation is bound by memory bandwidth, a big model on Strix Halo generates noticeably slower than the same model would on a discrete GPU that could hold it — and prompt processing, which is compute-heavy, lags further still. The exact same tradeoff shows up on the DGX Spark and, in milder form, on Apple Silicon: huge memory, moderate bandwidth, patience required.

So the honest framing is this. Strix Halo is the cheapest way to run a 70B-class model at home in a single quiet box — but it's not the fastest way to run one. If your priority is capacity and efficiency (big models, low power, tiny footprint), it's brilliant. If your priority is speed on models that already fit a 24GB card, a discrete GPU will run circles around it. Know which you're optimising for before you buy.

Verdict

Capacity champion, bandwidth compromise

Strix Halo's up-to-96GB VRAM allocation is a genuine breakthrough for capacity: it runs 70B-class models — even Mixtral 8x22B — in a silent, efficient mini-PC for well under a discrete multi-GPU build. The catch is bandwidth: ~256 GB/s means those big models run slower than they would on a discrete GPU that could hold them, and prompt processing especially lags. Buy it to run models bigger than a 24GB card allows, quietly and cheaply. Don't buy it expecting discrete-GPU speed.

Best for: Yes: 70B+ models at home, low power, tiny footprint, capacity over speed. No: maximum speed on models that already fit 24GB, heavy prompt-processing workloads.

The questions people actually ask

How much VRAM can the Ryzen AI Max+ 395 use?
Up to 96GB on a 128GB system, allocated via AMD Variable Graphics Memory in the BIOS, leaving about 32GB for the OS. On smaller-memory configurations (the chip ships in 32–128GB variants) the ceiling scales down accordingly. 96GB is the headline figure and it's what makes the platform interesting for running large models locally.
Is Strix Halo's VRAM real VRAM?
It's unified memory presented to the GPU as VRAM — the CPU, integrated GPU and system RAM share one physical pool, and you decide how much the GPU claims. Functionally, for loading and running models, it behaves like VRAM. The difference from a discrete GPU is bandwidth: this shared memory runs at ~256 GB/s versus 1,000+ on a dedicated card, which affects speed but not whether a model fits.
Can it really run Llama 70B?
Yes — that's the whole point. At 96GB allocated, it fits Llama 3.3 70B at Q6 (~58GB) with room for context, plus Qwen 72B and even Mixtral 8x22B. What it can't do is run them as fast as a discrete GPU that could hold them, because of the bandwidth gap. So it's the cheapest single box that runs 70B locally, at a speed that's usable rather than fast.
Strix Halo or a Mac Studio for local AI?
They're close cousins — both unified-memory machines that trade bandwidth for capacity. Strix Halo mini-PCs undercut a comparable Mac Studio on price and run Windows/Linux natively; a Mac Studio offers higher memory bandwidth on its top chips and the macOS/MLX ecosystem. For pure value on big-model capacity, Strix Halo; for bandwidth and the Apple software stack, the Mac. We compare the unified-memory-vs-discrete question in more depth separately.

The one-line takeaway: Strix Halo turns system RAM into AI memory, and up to 96GB of it, which is a genuinely new capability at this price. Just remember you're buying capacity, not speed. Size the models you want against that 96GB in our VRAM calculator, and if you're weighing it against a discrete build, our Mac vs NVIDIA and is-the-5090-worth-it pieces cover the speed side of the trade.

Hardware Editor

Ravi Malhotra

Ravi has been building and taking apart PCs since the single-core days — his idea of a good weekend is a repaste and a spreadsheet full of thermals. He covers GPUs, CPUs and the build decisions that actually move frame rates, and he'd rather hand you a benchmark than a press release.

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading