ExLlamaV2 is open-source and free (MIT), and on a GPU it's roughly twice as fast as llama.cpp. The catch: your model has to fit entirely in VRAM.…
The short answer
Aliteq
is ExLlama free? yes — and it's the fastest way to run a local model, with one big condition