
your GPU isn't too small for that 235B model — you're just using it wrong
A handful of llama.cpp flags let a 16GB card run models that shouldn't fit, and most people running local AI have never touched them.
Lena Fischer · 5d ago · 7 min
6 articles · newest first

A handful of llama.cpp flags let a 16GB card run models that shouldn't fit, and most people running local AI have never touched them.
Lena Fischer · 5d ago · 7 min

A new hardware breakdown puts real numbers on the two components most local-AI buying guides skip entirely — and skimping on either turns a snappy agent into a 10-minute wait.
Lena Fischer · Aug 19 · 7 min

the model that made headlines for being nearly free to rent turns out to be one of the most expensive things you could try to self-host
Lena Fischer · Aug 4 · 6 min

Phi-4 Mini, Gemma 3, and Llama 3.2 all advertise roughly the same 128K context window. Benchmark data says that number means something very different for each one.
Lena Fischer · Aug 4 · 7 min

Phi-4 Mini, Gemma 3, and Llama 3.2 all claim to fit — I checked which one actually leaves room to breathe.
Lena Fischer · Aug 4 · 6 min

Four real paths to running a 70B-parameter model on your own hardware or by the hour — priced out with actual 2026 numbers.
Lena Fischer · Aug 4 · 8 min