
I compared Ollama, vLLM and LM Studio — only one of them survives a second user
at one person typing, all three are basically tied. add a second and the gap turns into a cliff
Lena Fischer · Aug 22 · 7 min
9 articles · newest first

at one person typing, all three are basically tied. add a second and the gap turns into a cliff
Lena Fischer · Aug 22 · 7 min

Day-zero support for the new Qwen3.8 27B just landed on two AMD boxes — one needs a graphics card, one doesn't — and the tokens-per-second numbers are real.
Lena Fischer · Aug 18 · 6 min

Android Police ran a real local model on a real phone as a daily driver. it's not a GPU replacement — but it covers more than the GPU-shopping conversation usually admits.
Lena Fischer · Aug 16 · 6 min

An Android app called RikkaHub Agent turns a phone into a local-LLM-powered agent that can write code, run Linux commands, and compile software on its own — as long as you keep saying yes.
Priya Nair · Aug 9 · 8 min

A hobbyist just proved a genuine language model can run entirely on an $8 microcontroller — no cloud, no GPU, no internet connection required.
Ravi Malhotra · Aug 6 · 6 min

Phi-4 Mini, Gemma 3, and Llama 3.2 all advertise roughly the same 128K context window. Benchmark data says that number means something very different for each one.
Lena Fischer · Aug 4 · 7 min

Four real paths to running a 70B-parameter model on your own hardware or by the hour — priced out with actual 2026 numbers.
Lena Fischer · Aug 4 · 8 min

Llama is the open model that kicked off the whole local-AI movement, and running it yourself is genuinely one command. Here's how, and which size fits your GPU.
Lena Fischer · Aug 3 · 10 min

Mistral's models are famously efficient — they punch above their size, so you get a lot on modest hardware. Here's how to run Mistral locally in one command, by size.
Lena Fischer · Aug 3 · 10 min