
AI
your GPU isn't too small for that 235B model — you're just using it wrong
A handful of llama.cpp flags let a 16GB card run models that shouldn't fit, and most people running local AI have never touched them.
Lena Fischer · 6d ago · 7 min
2 articles · newest first

A handful of llama.cpp flags let a 16GB card run models that shouldn't fit, and most people running local AI have never touched them.
Lena Fischer · 6d ago · 7 min

AMD's own numbers put its workstation GPU at 53 tokens a second on Meta's new 30B model. Nvidia's flagship does more — for over three times the price.
Ravi Malhotra · Aug 11 · 6 min