
your GPU isn't too small for that 235B model — you're just using it wrong
A handful of llama.cpp flags let a 16GB card run models that shouldn't fit, and most people running local AI have never touched them.
Lena Fischer · Aug 25 · 7 min
4 articles · newest first

A handful of llama.cpp flags let a 16GB card run models that shouldn't fit, and most people running local AI have never touched them.
Lena Fischer · Aug 25 · 7 min

Nemotron 3.5 Lightning is a 30-billion-parameter model with only about 3 billion active per token — Nvidia says that's why it runs up to 4x faster than comparable open models on a single consumer GPU.
Ravi Malhotra · Aug 13 · 7 min

it beats Claude on some benchmarks and rivals GPT on others, and it's free to download next week. here's the hardware math that ruins the fun.
Ravi Malhotra · Aug 3 · 7 min

A '30B' model that runs as fast as a 3B one? That's a Mixture-of-Experts model, and it changes the local-AI math. Here's what MoE means and why it matters for what you can run.
Lena Fischer · Aug 3 · 10 min