how to run a local AI on a Raspberry Pi — a pocket-sized offline chatbot for $80

Yes, you can run a real language model on a Raspberry Pi 5. It won't be fast, but for a tiny, private, always-on AI it's genuinely fun and useful.…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short version

Pi 5 (8GB) runs 1-3B models via Ollama — usable for a small, offline, always-on AI.

The short version

Fastest: Gemma 3 1B at ~18-22 tok/s. Smartest: Phi-3 Mini (3.8B) at ~4-7 tok/s.

The short version

Add an active cooler + a USB SSD — the Pi throttles when hot, and an SSD helps with swap.

The short version

7B models technically load at aggressive quantization but run <1 tok/s — not practical.

The short version

The Hailo AI HAT+ doesn't help LLMs — it accelerates vision models, not language models.

The short version

Great for: offline assistant, edge AI, home automation, learning — not for heavy work.

Aliteq

Read the full story

how to run a local AI on a Raspberry Pi — a pocket-sized offline chatbot for $80

Read the full story on Aliteq